Models API

Multimodal models for real-world use

Process text, images, and documents, with computer-use capabilities for any digital environment.

Fast & Open

Holo3-1-35B-A3B

holo3-1-35b-a3b

Near-flagship accuracy at a fraction of the cost and latency. Fully open under Apache 2.0, on the free (10 RPM) and paid tiers.

Context
66KTokens
Input
$0.25$0.025 cached
Output
$1.80Per 1M tokens

MoE, 35B total, 3B active

Flagship

Holo3-122B-A10B

holo3-122b-a10b

Best-in-class reasoning and navigation for complex multi-step tasks across web, desktop, and mobile. Paid tier, with higher rate limits.

Context
66KTokens
Input
$0.40$0.04 cached
Output
$3.00Per 1M tokens

MoE, 122B total, 10B active

Rate-limited free access to Holo3-1-35B-A3B. No credit card required. Pay-as-you-go for every model, with higher rate limits.

After you pay, it can take up to 15 minutes for the system to process your payment and credit your balance.

Frequently asked questions

For maximum accuracy on complex, multi-step tasks, especially in novel environments, use Holo3-122B-A10B. For latency-sensitive workloads, cost-efficient automation, or well-defined tasks, Holo3-1-35B-A3B delivers Pareto-optimal accuracy at significantly lower cost. Both share the same Models API, so switching requires only a model ID change.

Multimodal models for real-world use

Process text, images, and documents, with computer-use capabilities for any digital environment. Rate-limited free tier, pay-as-you-go beyond.