New: SLMs outperform frontier LLMs on narrow enterprise tasks by up to 34% — Read the ScaleDown AI benchmark →
Products
⚡ Platform 🧠 Model Catalog 📖 Documentation
Solutions
🏥 Healthcare ⚖️ Legal 📈 Finance 🔧 Manufacturing
More
📊 Benchmarks 💳 Pricing 🔑 Sign In Deploy a model →
Pricing

Pay for what you
actually run.

No GPU reservations. No minimum commits. Inference billed per token. Training billed per GPU-hour. Cancel anytime.

Starter
Prototype
$0 / month

Deploy and test models free. Perfect for benchmarking SLMs before committing to production.

  • 1 model deployment
  • 1M tokens / month free
  • Shared GPU fleet
  • Community support
  • Public benchmark suite
Start free
Enterprise
Custom
Custom

Private VPC, compliance packages, unlimited deployments, and a dedicated ML engineering team.

  • Unlimited deployments
  • Private VPC / air-gapped options
  • HIPAA BAA + SOC 2 Type II
  • Custom fine-tuning pipelines
  • Dedicated ML engineering
  • 99.99% uptime SLA
  • Custom contracts
Talk to sales
Inference Pricing

Per-token rates by model class

All inference is billed per 1,000 tokens. No minimums. No rounding up.

ModelHardwareInput tokensOutput tokensp50 latencyIncluded in Production
SLM 1–3BA10G$0.00004/1k$0.00008/1k5–12ms
SLM 7–13BA10G$0.00012/1k$0.00022/1k18–35ms
LLM 30–70BH100$0.0008/1k$0.0016/1k80–160ms+add-on
LLM 405B+8×H100$0.004/1k$0.008/1k200–400msEnterprise only
FAQ

Common questions

Can I bring my own fine-tuned model?

Yes. Upload any HuggingFace-compatible checkpoint and deploy it in under 60 seconds. We support GGUF, GPTQ, AWQ, and full-precision formats.

Is my data used for training?

Never. Your inference data and training data are completely isolated and never used to improve any other customer's models or Quantum Junction's base models.

What compliance certifications do you have?

SOC 2 Type II, ISO 27001, and HIPAA BAA available on Production and Enterprise plans. FedRAMP Moderate in progress for Q3 2026.

How does smart routing work?

Our query complexity scorer evaluates each request and routes simple, narrow tasks to your SLM and complex, open-ended requests to your LLM — transparently, with full logs of every routing decision.

Ready to deploy your exact-fit model?

Start free in under 60 seconds. No GPU reservations. No minimum commits.

Deploy now — it's free → Read the docs