Qwen3.8-27B
Model Profile, Runtime Support, and FP8 VariantA documentation-based profile of Qwen3.8-27B, including its native vision-language capabilities, 262K context, thinking controls, supported runtimes, Apache 2.0 license, and official FP8 variant.
This page organizes first-party documentation and editorial analysis. ModelRun Lab has not independently reproduced the published performance or hardware claims.
Model overview
Qwen3.8-27B is a 27-billion-parameter dense causal language model with a vision encoder. Qwen describes it as a native vision-language model for text, image, and video understanding, with an emphasis on coding, professional work, research, and long-running agentic tasks.
This page summarizes the official Qwen model cards. ModelRun Lab has not independently run the model, measured its memory use, or reproduced Qwen's benchmark claims.
Published specifications
Qwen lists the following core specifications:
- 27B language-model parameters
- 64 layers
- Native context length of 262,144 tokens
- Extension up to one million tokens
- Native image and video understanding
- Thinking mode enabled by default
- Per-request thinking control through
reasoning_effort - Optional retention of reasoning context through
preserve_thinking
The hosted Qwen Cloud version is described as coming soon and is expected to provide a one-million-token context by default plus official built-in tools. Availability should be checked directly before planning around the hosted service.
Runtime compatibility
The official repository provides model weights and configuration files in Hugging Face Transformers format. Qwen lists compatibility with:
- Hugging Face Transformers
- vLLM
- SGLang
- TokenSpeed
Compatibility does not establish that every feature is supported identically across every runtime. Vision inputs, long context, thinking controls, and tool integrations should be checked against the runtime's current documentation.
Base weights and FP8 variant
Qwen publishes two closely related repositories:
Qwen/Qwen3.8-27B: the primary post-trained modelQwen/Qwen3.8-27B-FP8: an official fine-grained FP8 quantization of the same model
The FP8 model card identifies Qwen3.8-27B as its base model and uses a block size of 128. Qwen says its performance metrics are nearly identical to the original model. That statement is provider-reported and has not been independently verified by ModelRun Lab.
The FP8 repository is a deployment option, not a separate model topic. Hardware compatibility depends on the selected runtime, kernel support, GPU architecture, context length, and multimodal workload.
Hardware boundary
The official model cards reviewed for this page do not provide one universal minimum consumer-GPU VRAM figure. A 27B model's usable memory footprint varies with precision, quantization, context length, KV cache, batch size, vision inputs, and runtime overhead.
For that reason, this page does not label a specific desktop GPU as an official minimum. Before deployment, consult the runtime documentation and test the exact precision and context configuration on the intended hardware.
Thinking controls
Qwen states that thinking mode is enabled by default. Applications can disable it per request, adjust reasoning depth with reasoning_effort, and preserve reasoning context from earlier messages with preserve_thinking.
These controls can affect latency and token use. Production systems should evaluate them with representative prompts rather than assuming the highest reasoning setting is always the best operational choice.
License
Both official Hugging Face repositories identify the license as Apache 2.0. This page reports the repository metadata and does not provide legal advice. Deployers remain responsible for reviewing the complete license, model terms, dependencies, and the rights attached to their inputs and outputs.
Evidence boundaries
- Architecture, context, features, runtime names, and license metadata come from Qwen's official model cards.
- Benchmark results and the claim of near-identical FP8 performance are provider-reported.
- No universal hardware minimum is stated because the official cards do not provide one.
- ModelRun Lab has not independently tested the model or its quantized variant.