Skip to main content
SambaStack is a turnkey AI inference platform delivering industry-leading performance, energy efficiency, and out-of-the-box support for popular open-source models. Available as both cloud (hosted) and on-premises solutions. With SambaStack, you can:
  • Deploy leading AI models – Access models from 9+ providers, including Meta, DeepSeek, Mistral AI, and Google, through a unified platform.
  • Reduce latency and inference costs – Use built-in prompt caching to achieve 90%+ cache hit rates under sustained workloads.
  • Deploy your way – Run on your infrastructure or use SambaNova’s fully managed service.

Who this guide is for

System administrators managing:
  • Hardware infrastructure
  • Kubernetes clusters
  • Inference services (models, user groups, access control)

Required skills

  • Linux system administration
  • Kubernetes operations (kubectl, Helm)
  • Log analysis and troubleshooting
  • System credential management
See the SambaStack release notes for new features, improvements, bug fixes, and version-specific updates. For inference service features and API usage, see the Developer Guide.

Deployment options

SambaStack hosted

A fully managed cloud service from SambaNova. Build and operate high-performance inference services without managing the underlying infrastructure – SambaNova provides and maintains the AI hardware and Kubernetes environment, while you configure and manage your AI services (model selection, user groups, and users).Quickstart guide

SambaStack on-prem

Runs in your own data center, giving you full control over the AI infrastructure, Kubernetes environment, and AI services you deploy. You manage everything end-to-end – model selection, user groups, and users – to support your organization’s needs.Quickstart guide

Deployment comparison

See the on-prem architecture diagram and connection flows for a full breakdown of how requests, administration, and supporting infrastructure connect.

Explore key capabilities

Supported models and bundles

Browse available models, context lengths, and features.

Deploying model bundles

Deploy your first model bundle on SambaStack.

Prompt caching

Cut latency and cost by caching repeated context – hit rates reach 90%+ under sustained traffic.

Authentication

Set up OIDC or LDAP authentication for your deployment.