dots3-note Preview: Hardware, FP8 Serving and 512K Context Explained
dots3-note Preview: Hardware, FP8 Serving and 512K Context Explained
dots3-note Preview is an unusually large open-weight multimodal agent model: 280 billion total parameters, 16 billion activated parameters per token, and a maximum 512K-token context window. That low active-parameter count can reduce compute per generated token, but it does not turn a 280B checkpoint into a 16B model for memory planning.
The official deployment guidance makes the practical boundary clear. Dots Studio recommends serving the FP8 checkpoint on one eight-GPU node with SGLang or vLLM. Its published vLLM example uses eight NVIDIA H100 GPUs, while BF16 requires still more memory. This is therefore a server-class model today, not a conventional single-consumer-GPU download.
Quick answer
| Question | Practical answer |
|---|---|
| Is it open weight? | Yes. Model weights and repository assets are released under Apache 2.0. |
| Model size | 280B total parameters, 16B activated |
| Context | Up to 512K tokens |
| Inputs | Text, image, video and audio |
| Output | Text |
| Checkpoints | BF16 and FP8 |
| Official serving direction | SGLang or vLLM, with an eight-GPU FP8 reference deployment |
| Single 24/32 GB consumer GPU? | Not a realistic official deployment target for the released full checkpoint |
| Tool calling | Supported in the documented SGLang/vLLM serving paths |
| Production maturity | Preview release; some framework integrations are still landing upstream |
The useful way to evaluate dots3-note is therefore not “does 16B active mean it fits like a 16B model?” It is: does its sparse architecture provide enough agentic and multimodal value to justify a large distributed-memory serving footprint?
1. Why 16B active parameters do not mean 16B-model VRAM
Mixture-of-Experts models route each token through only a subset of their experts. dots3-note has 256 routed experts plus one shared expert and selects eight routed experts for a token. That is how a 280B-total model can activate roughly 16B parameters during computation.
But the model still needs its expert weights available to the serving system. Active parameters primarily describe compute sparsity, not the total checkpoint storage requirement.
A rough weight-only thought experiment illustrates the distinction:
- 280B parameters at 2 bytes each would be roughly 560 GB before runtime overhead if represented uniformly as BF16;
- FP8 can roughly halve the bytes-per-weight component relative to BF16, but the complete serving footprint also includes non-weight state, runtime buffers, multimodal components and KV cache;
- actual checkpoint and runtime memory should be taken from the released artifacts and serving stack, not estimated from “16B active” alone.
This is why the official recipe targets distributed serving rather than a normal desktop GPU.
2. What Dots Studio actually released
The official repository describes dots3-note Preview as the first open-weight model in the dots3 family and the family's most lightweight member. Its published architecture includes:
| Component | Official specification |
|---|---|
| Language architecture | Multimodal MoE |
| Total parameters | 280B |
| Activated parameters | 16B |
| Transformer depth | 1 dense + 45 MoE layers |
| Experts | 256 routed + 1 shared, top-8 routing |
| Context | 512K tokens |
| Vision encoder | MoE ViT, 7B total / 1.2B activated |
| Audio encoder | Dense, 800M |
| Precision | BF16, FP8 |
| Modalities | Text, image, video, audio input; text output |
The model is positioned for general reasoning, coding, tool use, multimodal understanding, long-context processing and multi-step agent workflows involving exploration and memory updates.
Those are vendor-described capabilities. Benchmark tables should be treated as release evidence, not as independent proof that the model will outperform alternatives in a particular deployment.
3. The realistic hardware envelope
Official reference: eight GPUs
Dots Studio recommends the FP8 checkpoint on one eight-GPU node. Its current vLLM recipe specifically demonstrates 8× H100 with tensor parallelism of eight plus expert parallelism.
The SGLang recipe similarly uses eight-way data, tensor and expert parallel settings. This matters more than a speculative desktop VRAM table: it tells operators what configuration the model's maintainers are actually exercising.
Why a huge context window raises memory pressure
The advertised 512K context is a model capability ceiling, not a sensible default for every deployment. Long context consumes additional KV-cache memory, and multimodal inputs can add their own processing and memory costs.
The official repository explicitly tells operators to tune context length according to available memory, concurrency and input modalities. The published vLLM example sets --max-model-len 262144, or 256K, rather than automatically exposing the full 512K maximum.
For real deployments, start with the smallest context window that satisfies the workload and increase it only after measuring memory headroom and latency.
Consumer GPUs
There is no sound basis for calling dots3-note a normal 16 GB, 24 GB or 32 GB GPU model merely because only 16B parameters are active during a token.
Community quantizations or aggressive offload schemes may eventually change the experimentation envelope, but they should not be confused with the maintainers' current supported deployment path. A future GGUF or lower-bit community build also would not automatically preserve the same throughput, multimodal support or 512K-context behavior.
4. FP8 versus BF16
Dots Studio publishes both BF16 and FP8 checkpoints. For this model, FP8 is the practical starting point because the official deployment recommendation itself is built around it.
FP8 advantages: lower weight-memory pressure and a more realistic distributed serving footprint on hardware with strong FP8 support.
BF16 advantages: a conventional higher-precision representation and potentially simpler reasoning about numerical behavior, at substantially higher memory cost.
The repository explicitly says BF16 requires more memory. It does not provide evidence for a universal quality delta between the two checkpoints, so operators should avoid assuming either “FP8 is identical” or “BF16 is always materially better” without workload-specific evaluation.
5. vLLM support is currently the cleanest documented path
The repository says native dots3-note Preview support is available on vLLM main, with a recent nightly build recommended until that support reaches a stable release.
Its reference command uses:
- eight-way tensor parallelism;
- expert parallelism;
- an FP8 checkpoint;
- a 256K maximum model length;
- optional multi-token-prediction speculative decoding;
- optional automatic tool calling with the
dotsparser.
That makes vLLM the clearest current choice for operators already running distributed OpenAI-compatible inference infrastructure.
The important maturity caveat is the word main: native support existing upstream does not mean every stable vLLM package already contains it. Pin and document the runtime revision used for a production evaluation.
6. SGLang has a dedicated release image and detailed recipe
Dots Studio also documents SGLang and points users to a dedicated lmsysorg/sglang:dev-dots3-note image. Its example exposes an OpenAI-compatible server and enables multimodal processing across an eight-GPU configuration.
The documented path can also enable automatic tool calling with --tool-call-parser dots.
One useful optimization is the model's MTP/NEXTN speculative path. The repository says this optional configuration can reduce time per output token by more than 50% in its serving setup. That is a maintainer-reported serving result, not a guarantee for every prompt mix or hardware configuration.
SGLang support is still being integrated upstream, so operators should treat the provided image/recipe as a more precise compatibility target than an arbitrary older SGLang installation.
7. Transformers support is still landing
The official repository currently points to a specific Hugging Face Transformers pull request rather than claiming ordinary stable-package support. It similarly notes an SGLang pull request still under review.
That is an important preview-status signal. A model can have released weights and working reference recipes while its ecosystem support is still converging.
For experimentation, following the pinned upstream revision can be reasonable. For a production service, waiting for integrations to reach stable releases reduces dependency drift and makes rebuilding the environment easier.
8. Multimodal support is broader than a typical text MoE
The model accepts:
- text;
- images;
- video;
- audio.
It produces text output. The official examples show image, audio and video URL inputs, and note that video can include its audio track when available.
This breadth matters when comparing dots3-note with text-only sparse agent models. The tradeoff is that multimodal components also make deployment and memory planning more complex than the language-model parameter count alone suggests.
For a text-only workload, the documented serving stacks can load only the language component. That may be useful when the application needs the reasoning/agent model but not its vision or audio path.
9. Tool calling and agent workloads
The release is explicitly aimed at multi-step workflows, including tool use, exploration, memory updates and adaptation. Both SGLang and vLLM documentation in the repository expose OpenAI-compatible tool-calling options.
That gives dots3-note a practical path into existing agent systems that already speak OpenAI-style APIs. But tool-call syntax support is not the same as reliable autonomous behavior.
Before adopting it for a long-running agent, test at least:
- tool-selection accuracy;
- argument/schema validity;
- recovery after tool failures;
- behavior over long histories;
- memory-update quality;
- latency and cost under repeated tool loops;
- whether multimodal inputs materially improve the target workflow.
Preview models especially need this application-level evaluation because release benchmarks cannot reproduce an organization's exact tools and failure modes.
10. What TEMPO means, and what it does not prove
Dots Studio describes the model as emphasizing long-horizon agent behavior and introduces TEMPO, a reinforcement-learning approach centered on test-time-scaled value estimation and macro-step policy optimization.
The interesting design goal is self-critique across long trajectories: an agent should be able to evaluate intermediate progress rather than only optimize a single short response.
That is relevant to research and agent-system design, but it should not be translated into a claim that the preview model is automatically reliable over extremely long autonomous runs. The full technical report is still listed as forthcoming in the repository.
Treat TEMPO as an architectural/training direction to evaluate, not a substitute for operational testing.
11. dots3-note versus smaller local agent models
The model occupies a different deployment class from 20B-30B dense or low-total-parameter models that can be quantized onto one high-memory consumer GPU.
Choose dots3-note Preview when:
- multimodal input is central;
- very long context is genuinely useful;
- an eight-GPU or comparable distributed inference environment is available;
- agent/tool workflows justify testing a large sparse model;
- preview-stage runtime pinning is acceptable.
Prefer a smaller model when:
- the service must run on one consumer GPU or a Mac;
- deployment simplicity matters more than maximum context;
- the workload is mostly text and short-context;
- stable packaged runtime support is required;
- power, hardware cost or operational complexity dominates the decision.
A 280B-total MoE can be compute-efficient relative to a similarly sized dense model while still being far less convenient to host than a genuinely small checkpoint.
12. Recommended evaluation plan
For teams with suitable hardware, the lowest-risk evaluation path is:
- Start with FP8. It is the maintainers' recommended deployment checkpoint.
- Use the documented vLLM or SGLang path. Avoid inventing a different stack before establishing a known-good baseline.
- Reduce context initially. Do not begin at 512K unless the workload requires it.
- Test text-only first, then add image/audio/video inputs if the application needs them.
- Measure tool-call reliability, not just benchmark-style question answering.
- Record exact runtime revisions. Some integrations are on main branches or pending pull requests.
- Evaluate concurrency and KV-cache pressure before promising capacity to users.
- Revisit the stack after stable releases land. Preview-specific pins should not become invisible infrastructure debt.
Bottom line
dots3-note Preview is genuinely open weight and technically distinctive: 280B total parameters, 16B activated, multimodal input, a 512K context ceiling, FP8/BF16 checkpoints and explicit agent/tool-serving paths.
But 16B active is not a 16B-model memory requirement. The maintainers' own FP8 recommendation is an eight-GPU node, and the vLLM reference uses eight H100s. That is the most important deployment fact for anyone deciding whether to download the model.
For operators with distributed GPU infrastructure, dots3-note is a credible preview to evaluate for multimodal, long-context agent workloads. For single-GPU local AI, smaller models remain the practical choice until materially smaller official or well-validated deployment artifacts change the equation.
Sources
- Dots Studio, dots3-note Preview official repository and deployment documentation: https://github.com/studio-dots-ai/dots3-note-prev
- Dots Studio technical blog: https://studio.dots.ai/dots/dots3-en.html
- Official model checkpoint: https://huggingface.co/dots-studio/dots3-note-prev
- Official FP8 checkpoint: https://huggingface.co/dots-studio/dots3-note-prev-fp8