Deployment options
When deploying models in SambaStack, administrators can select from various context length and batch size combinations.- Smaller batch sizes provide higher token throughput (tokens/second).
- Larger batch sizes provide better concurrency for multiple users.
Finding models on your cluster
You can run the following command to discover available models in your cluster:kubectl get models does not return the name you send to the API
kubectl get models does not return the name you send to the API
The command lists Kubernetes resource names (The two often differ only in case or punctuation, so the mismatch is easy to miss. In the example above they differ by a single character.To read the serving name of a deployed model, run:A model can also define
metadata.name). Inference requests must use the serving name (spec.name), which is what the Model ID headings below use.spec.aliases, additional names that route to the same model. For the full Model field reference, see Deploy custom checkpoints.Text to speech is one case where the names diverge: address /v1/audio/speech as qwen3-tts, not by the talker or vocoder resource names. See Text to speech.All supported models
Search or filter to find a model, then select it to jump to its full configuration.Meta
Meta-Llama-3.3-70B-Instruct
Meta-Llama-3.1-8B-Instruct
Meta-Llama-3.1-70B-Instruct
Meta-Llama-3.1-405B-Instruct
Llama-4-Maverick-17B-128E-Instruct
MiniMax
MiniMax-M3
MiniMax-M2.7
MiniMax-M2.5
Mistral AI
Mistral-Large-3-675B-Instruct-2512
DeepSeek
DeepSeek-R1-0528
DeepSeek-R1-Distill-Llama-70B
DeepSeek-V3-0324
DeepSeek-V3.1
DeepSeek-V3.2
DeepSeek-V3.1-Terminus
OpenAI
gpt-oss-120b
gpt-oss-20b
Whisper-Large-v3
gemma-3-27b-it
gemma-3-12b-it
gemma-4-31B-it
Alibaba Cloud
Qwen3-235B-A22B-Instruct-2507
Qwen3-32B
Qwen3-TTS-Talker
Qwen3-TTS-Vocoder
Tokyotech-llm
Llama-3.3-Swallow-70B-Instruct-v0.4
Other
E5-Mistral-7B-Instruct
Recommended model bundles
In SambaStack, a bundle is a packaged deployment that groups one or more models together with their associated configurations, such as batch size and sequence length. A single model can also be deployed on its own by pairing it with a model profile, without creating a bundle. For example, deploying theMeta‑Llama‑3.3‑70B model with a batch size of 4 and a sequence length of 16K tokens constitutes a single configuration. A bundle, however, can contain multiple such configurations, either for the same model or for different models.
SambaNova’s RDU technology enables several models and configurations to be loaded simultaneously in a single deployment. This allows you to switch instantly between models and between batch‑/sequence‑size profiles as needed. In contrast to traditional GPU systems, where deployments are typically single‑model and static, SambaStack supports multi‑model, multi‑configuration bundles. This approach delivers higher efficiency, greater flexibility, and increased throughput while preserving low latency.
You can run the following command to discover available bundles in your cluster:
If the bundles listed below do not satisfy your inference requirements, you can create custom bundles that combine any mix of models and configurations so long as they fit in DDR memory.
Suggested bundles per model
For each model, this section lists the suggested bundle for typical use and any alternative bundles that trade off context length, batch size, or modality support. See Bundle configurations below for the seq length/batch size details of each bundle.Meta
Meta-Llama-3.3-70B-InstructSuggested: 70b-3dot3-ss-4-8-16-32-64-128kAlternatives: 70b-3dot3-ss-full-whisper, us-agentic-rag-1-1, e5-mistral-70b-64k-128kMeta-Llama-3.1-8B-InstructSuggested: us-agentic-rag-1-1Alternatives: qwen3-32b-llama405b-s-mMeta-Llama-3.1-405B-InstructSuggested: qwen3-32b-llama405b-s-mLlama-4-Maverick-17B-128E-InstructSuggested: llama-4-medium-8-16-32-64-128kAlternatives: llama-4-medium-ss-16k-bs24, us-agentic-rag-1-1MiniMax
MiniMax-M3PREVIEWSuggested: minimax-m3-32-64-128-256-512k-1m (full 32K–1M context range)Alternatives: minimax-m3-32kMiniMax-M2.7Suggested: dyt-minimax-m2p7-32-64-160-192k-pc (adds prompt caching)Alternatives: dyt-minimax-m2p7-32k-pc, dyt-minimax-m2p7-32-64-160-192k, dyt-minimax-m2p7-32-160-192k, dyt-minimax-m2p7-32k-v2MiniMax-M2.5Suggested: dyt-minimax-m2p5-32-160kAlternatives: dyt-minimax-m2p5-32kMistral AI
Mistral-Large-3-675B-Instruct-2512PREVIEWSuggested: mistral-large-3-fp8-8-16-32kAlternatives: mistral-large-3-fp8-8kDeepSeek
DeepSeek-R1-0528Suggested:deepseek-4in1-fp8-128k(higher context length)
deepseek-r1-v3-fp8-8k, deepseek-r1-v31-fp8-8kDeepSeek-V3-0324Suggested:deepseek-4in1-fp8-128k(higher context length)
deepseek-r1-v3-fp8-8k, deepseek-v3-v31-fp8-8k, deepseek-v3-v3termi-fp8-8kDeepSeek-V3.1Suggested:deepseek-4in1-fp8-128k(higher context length)
deepseek-r1-v31-fp8-8k, deepseek-v3-v31-fp8-8kDeepSeek-V3.1-TerminusSuggested:deepseek-4in1-fp8-128k(higher context length)
deepseek-v3-v3termi-fp8-8kOpenAI
gpt-oss-120bSuggested: us-agentic-rag-1-1Alternatives: cd-dyt-gpt-oss-120b-8-32-64-128k, gpt-gemma-whisper-mistralgpt-oss-20bPREVIEWSuggested: dyt-gpt-oss-20b-32-64-128kWhisper-Large-v3Suggested: qwen3-32b-whisper-e5-mistralAlternatives: 70b-3dot3-ss-full-whisper (known issue: does not load within default startup time), gpt-gemma-whisper-mistralgemma-3-27b-itPREVIEWSuggested: gemma3-27b-32-128kgemma-3-12b-itPREVIEWSuggested: gemma3-v3Alternatives: gpt-gemma-whisper-mistralgemma-4-31B-itSuggested:gemma-4-31b-32-128-256k(adds 256K context; text, image, and video support)gemma-4-31b-32-128k(up to 128K context; text, image, and video support, higher throughput)
Alibaba Cloud
Qwen3-235B-A22B-Instruct-2507Suggested: dyt-qwen3-235b-32-128kQwen3-32BSuggested: qwen3-32b-whisper-e5-mistralAlternatives: qwen3-32b-llama405b-s-mQwen3-TTS-TalkerPREVIEWSuggested: qwen3-tts-talkerQwen3-TTS-VocoderPREVIEWSuggested: qwen3-tts-vocoderOther
E5-Mistral-7B-InstructSuggested: us-agentic-rag-1-1Alternatives: e5-mistral-70b-64k-128k, qwen3-32b-whisper-e5-mistral, gpt-gemma-whisper-mistralBundle configurations
The table below lists the configuration details for each bundle referenced above.
†
cd-dyt-gpt-oss-120b-8-32-64-128k supports sequence lengths from 8K–128K, and dyt-gpt-oss-20b-32-64-128k supports 32K–128K. If you require shorter context lengths (4K–16K) for either model, contact your SambaNova representative.
