Models API
Multimodal models for real-world use
Process text, images, and documents, with computer-use capabilities for any digital environment.
Holo3-1-35B-A3B
holo3-1-35b-a3b
Near-flagship accuracy at a fraction of the cost and latency. Fully open under Apache 2.0, on the free (10 RPM) and paid tiers.
- Context
- 66KTokens
- Input
- $0.25$0.025 cached
- Output
- $1.80Per 1M tokens
MoE, 35B total, 3B active
Holo3-122B-A10B
holo3-122b-a10b
Best-in-class reasoning and navigation for complex multi-step tasks across web, desktop, and mobile. Paid tier, with higher rate limits.
- Context
- 66KTokens
- Input
- $0.40$0.04 cached
- Output
- $3.00Per 1M tokens
MoE, 122B total, 10B active
Rate-limited free access to Holo3-1-35B-A3B. No credit card required. Pay-as-you-go for every model, with higher rate limits.
After you pay, it can take up to 15 minutes for the system to process your payment and credit your balance.
Tutorials to help you deploy our models in minutes
Frequently asked questions
For maximum accuracy on complex, multi-step tasks, especially in novel environments, use Holo3-122B-A10B. For latency-sensitive workloads, cost-efficient automation, or well-defined tasks, Holo3-1-35B-A3B delivers Pareto-optimal accuracy at significantly lower cost. Both share the same Models API, so switching requires only a model ID change.
Multimodal models for real-world use
Process text, images, and documents, with computer-use capabilities for any digital environment. Rate-limited free tier, pay-as-you-go beyond.