No GPU reservations. No minimum commits. Inference billed per token. Training billed per GPU-hour. Cancel anytime.
Deploy and test models free. Perfect for benchmarking SLMs before committing to production.
For teams running SLMs in production. Dedicated GPU capacity, private endpoints, and SLA guarantees.
Private VPC, compliance packages, unlimited deployments, and a dedicated ML engineering team.
All inference is billed per 1,000 tokens. No minimums. No rounding up.
| Model | Hardware | Input tokens | Output tokens | p50 latency | Included in Production |
|---|---|---|---|---|---|
| SLM 1–3B | A10G | $0.00004/1k | $0.00008/1k | 5–12ms | ✓ |
| SLM 7–13B | A10G | $0.00012/1k | $0.00022/1k | 18–35ms | ✓ |
| LLM 30–70B | H100 | $0.0008/1k | $0.0016/1k | 80–160ms | +add-on |
| LLM 405B+ | 8×H100 | $0.004/1k | $0.008/1k | 200–400ms | Enterprise only |
Yes. Upload any HuggingFace-compatible checkpoint and deploy it in under 60 seconds. We support GGUF, GPTQ, AWQ, and full-precision formats.
Never. Your inference data and training data are completely isolated and never used to improve any other customer's models or Quantum Junction's base models.
SOC 2 Type II, ISO 27001, and HIPAA BAA available on Production and Enterprise plans. FedRAMP Moderate in progress for Q3 2026.
Our query complexity scorer evaluates each request and routes simple, narrow tasks to your SLM and complex, open-ended requests to your LLM — transparently, with full logs of every routing decision.
Start free in under 60 seconds. No GPU reservations. No minimum commits.