Research
How models understand and act on software.
The H Research Lab develops vision-language and action models for computer use. We study visual grounding, action generation, long-horizon control, and evaluation across web, desktop, and mobile.
What we study
Computer use is a research problem.
A useful computer-use model has to understand an interface, choose an action, observe what changed, and keep working until the task is complete.
We post-train models for perception and action. We also build the interactive environments those models learn in, and evaluations that measure whether a task was completed.
01
Visual grounding
Teaching models to connect language with the exact controls, content, and state visible in a software interface. For example, translating "Click OK" into a click at (500, 600).
02
Action generation
Given a task goal, deciding what to do next: click, type, scroll, call a tool, or stop. The screen is the model's source of context.
03
Long-horizon control
Planning multi-step tasks, checking progress, and recovering from errors and unexpected situations.
04
Evaluation and data
Building interactive environments, training data, and benchmarks that measure whether an agent completed the work correctly, so results can be compared with other systems.
How we train
Environments built for experimentation.
Interactive environments
Synthetic software environments let us vary tasks, interface states, and failure cases without using customer data.
Post-training and evaluation
We train models on action sequences, then test whether they can complete unseen tasks and recover from mistakes.
Open models
Research released through the Holo family.
We publish model weights, technical reports, and evaluations so other teams can study and build on our work.
The team
A team with academic and applied research experience.
The Research Lab brings together academic researchers, applied scientists, and engineers with experience in model training and deployed agent systems.
Their previous work covers reinforcement learning, alignment, multi-agent systems, applied mathematics, theoretical physics, computational neuroscience, and computer use.
Previous research environments
Team members previously studied, taught, or conducted research at these institutions.
- University of California
- Flatiron Institute
- Mila, Quebec AI Institute
- ServiceNow Research
- École polytechnique / IASD
- ENS Paris-Saclay
Previous industry teams
Team members have also worked in model labs and production AI teams.
- AI21 Labs
- Cohere
- DeepMind
- Huawei
- InstaDeep
- DeepFlow
Areas represented in the team
- Reinforcement learning
- Alignment
- Multi-agent systems
- Applied mathematics
- Theoretical physics
- Computational neuroscience
- Computer use
How we work
From research question to tested system.
Researchers work with the engineers who build training environments, run evaluations, and deploy models. Product use gives the lab concrete failure cases to investigate. New methods return to the product only after they hold up in evaluation.
Continue from here
Work with the Research Lab
We hire researchers and engineers to work on vision-language models, interactive training environments, post-training, and evaluation for computer-use agents.