Models API
Multimodal models for real-world use
Process text, images, and documents, with computer-use capabilities for any digital environment.
Holo4-35B-A3B
holo4-35b-a3b
Next-generation MoE with a longer context window. Near-flagship accuracy at low cost and latency, built for computer-use.
- OSWorld
- 80.8%Desktop tasks
- Context
- 262KTokens
- Input
- $0.30$0.03 cached
- Output
- $2.00Per 1M tokens
MoE, 35B total, 3B active
Holo4-27B
holo4-27b
Dense flagship-class reasoning for complex multi-step tasks across web, desktop, and mobile, with an extended context window.
- OSWorld
- 85.2%Desktop tasks
- Context
- 262KTokens
- Input
- $0.40$0.04 cached
- Output
- $3.00Per 1M tokens
Dense, 27B
Previous generation, still served
| Model | Context | Input | Output |
|---|---|---|---|
Holo3-1-35B-A3B holo3-1-35b-a3b | 65,536 | $0.25 | $1.80 |
Holo3-122B-A10B holo3-122b-a10b | 65,536 | $0.40 | $3.00 |
Rate-limited free access to Holo3-1-35B-A3B. No credit card required. Pay-as-you-go for every model, with higher rate limits.
After you pay, it can take up to 15 minutes for the system to process your payment and credit your balance.
Tutorials to help you deploy our models in minutes
Frequently asked questions
Start with Holo4. Holo4-35B-A3B is the default for production computer-use workloads: longer context, strong accuracy, and low cost and latency. For denser flagship-class reasoning on complex multi-step tasks, use Holo4-27B. Earlier Holo3 models remain available on the same Models API if you need them. Switching only requires a model ID change.
Multimodal models for real-world use
Process text, images, and documents, with computer-use capabilities for any digital environment. Rate-limited free tier, pay-as-you-go beyond.