Publications

Papers, open models, datasets, and engineering notes from the Research Lab.

Everything the lab has published, with the primary artifact attached: arXiv for papers, Hugging Face for weights and benchmarks, and the blog for the systems work behind them.

Papers, open models, datasets, and engineering notes

Filter by type or research area. Every entry links to its primary artifact.

Type

Research area

9 of 17 entries

NeoMME

2026 · Open model · Apache 2.0

Multimodal encoders and retrievers for long-context document understanding, released at 260M and 800M parameters with dense and late interaction retriever variants.

Multimodal retrieval

DragOn: a drag-grounding benchmark and training dataset for GUI agents

2026 · Dataset and benchmark · Apache 2.0

Screenshots paired with an instruction and start and end bounding boxes, so models can be trained and measured on drag and drop, an action most grounding datasets leave out.

Evaluation and dataVisual grounding

Holo3.1

2026 · Open model · Apache 2.0

Open-weight action models at 0.8B, 4B, 9B, and 35B-A3B, with quantized checkpoints for local and single-GPU deployment.

Visual groundingAgents and action

Holotron 3 Nano

2026 · Open model · NVIDIA Open Model License

A computer-use model trained on NVIDIA's Nemotron 3 Nano Omni backbone, part of H's work on efficient open backbones.

Agents and actionVisual grounding

Holo3

2026 · Open model · Apache 2.0

The third generation of Holo action models, trained for grounding and multi-step control on web and desktop interfaces.

Visual groundingAgents and actionLong-horizon control

Holotron-12B

2026 · Open model · NVIDIA Open Model License

A 12B computer-use model built on NVIDIA Nemotron Nano V2 VL, released as part of H's work inside the Nemotron ecosystem.

Visual groundingAgents and action

Holo2-235B-A22B preview

2026 · Open model · CC BY-NC 4.0

A research preview of the largest Holo2 checkpoint, a mixture of experts model with 235B total and 22B active parameters.

Agents and actionLong-horizon control

Holo2

2025 · Open model · CC BY-NC 4.0

Holo2 at 4B, 8B, and 30B-A3B, trained for grounding and action prediction across web, desktop, and mobile screens.

Visual groundingAgents and action

Surfer 2: The Next Generation of Cross-Platform Computer Use Agents

2025 · Methods paper · Preprint · 53 authors, listed alphabetically

A unified agent architecture that works purely from visual observations, so a single system operates web, desktop, and mobile environments without environment-specific interfaces.

Agents and actionLong-horizon controlEvaluation and data