- Deploy leading AI models – Access models from 9+ providers, including Meta, DeepSeek, Mistral AI, and Google, through a unified platform.
- Reduce latency and inference costs – Use built-in prompt caching to achieve 90%+ cache hit rates under sustained workloads.
- Deploy your way – Run on your infrastructure or use SambaNova’s fully managed service.
Who this guide is for
System administrators managing:
- Hardware infrastructure
- Kubernetes clusters
- Inference services (models, user groups, access control)
Required skills
- Linux system administration
- Kubernetes operations (
kubectl, Helm) - Log analysis and troubleshooting
- System credential management
See the SambaStack release notes for new features, improvements, bug fixes, and version-specific updates. For inference service features and API usage, see the Developer Guide.
Deployment options
SambaStack hosted
A fully managed cloud service from SambaNova. Build and operate high-performance inference services without managing the underlying infrastructure – SambaNova provides and maintains the AI hardware and Kubernetes environment, while you configure and manage your AI services (model selection, user groups, and users).Quickstart guide
SambaStack on-prem
Runs in your own data center, giving you full control over the AI infrastructure, Kubernetes environment, and AI services you deploy. You manage everything end-to-end – model selection, user groups, and users – to support your organization’s needs.Quickstart guide
Deployment comparison
See the on-prem architecture diagram and connection flows for a full breakdown of how requests, administration, and supporting infrastructure connect.
Explore key capabilities
Supported models and bundles
Browse available models, context lengths, and features.
Deploying model bundles
Deploy your first model bundle on SambaStack.
Prompt caching
Cut latency and cost by caching repeated context – hit rates reach 90%+ under sustained traffic.
Authentication
Set up OIDC or LDAP authentication for your deployment.

