Skip to main content
SambaStack is a turnkey AI inference platform delivering industry-leading performance, energy efficiency, and out-of-the-box support for popular open-source models. Available as both cloud (hosted) and on-premises solutions. With SambaStack, you can:
  • Deploy leading AI models – Access models from 9+ providers, including Meta, DeepSeek, Mistral AI, and Google, through a unified platform.
  • Reduce latency and inference costs – Use built-in prompt caching to achieve 90%+ cache hit rates under sustained workloads.
  • Deploy your way – Run on your infrastructure or use SambaNova’s fully managed service.

Who this guide is for

System administrators managing:
  • Hardware infrastructure
  • Kubernetes clusters
  • Inference services (models, user groups, access control)

Required skills

  • Linux system administration
  • Kubernetes operations (kubectl, Helm)
  • Log analysis and troubleshooting
  • System credential management
See the SambaStack release notes for new features, improvements, bug fixes, and version-specific updates. For inference service features and API usage, see the Developer Guide.

Deployment options

SambaStack hosted

A fully managed cloud service from SambaNova. Build and operate high-performance inference services without managing the underlying infrastructure – SambaNova provides and maintains the AI hardware and Kubernetes environment, while you configure and manage your AI services (model selection, user groups, and users).Quickstart guide

SambaStack on-prem

Runs in your own data center, giving you full control over the AI infrastructure, Kubernetes environment, and AI services you deploy. You manage everything end-to-end – model selection, user groups, and users – to support your organization’s needs.Quickstart guide

Deployment comparison

See the on-prem architecture diagram and connection flows for a full breakdown of how requests, administration, and supporting infrastructure connect.

Explore key capabilities

Supported models and bundles

Browse available models, context lengths, and features.

Deploy a model bundle

Deploy your first model bundle on SambaStack.

Prompt caching

Cut latency and cost by caching repeated context – hit rates reach 90%+ under sustained traffic.

Authentication

Set up OIDC or LDAP authentication for your deployment.