Researchers at Peking University have developed UNIBROWSE, a comprehensive training framework designed to improve AI agents' ability to navigate and interact with multimodal web content.

The framework addresses a critical gap in existing agent training methods, which typically handle only text-based and image-to-text browsing patterns while ignoring text-to-image interactions.

UNIBROWSE introduces what the researchers call a "unified data pipeline" that generates training data covering all three information-flow patterns found in real-world web browsing. The system augments curated knowledge graphs with live web retrieval and uses a novel "exploration degree" metric to filter low-quality training instances.

Training methodology and performance

The team trained a 35-billion parameter agent using supervised fine-tuning combined with exploration-aware reinforcement learning on the generated data. The resulting model achieved 54.4% average accuracy across five multimodal browsing benchmarks.

This represents a 10.5 percentage point improvement over the base Qwen3.5-35B-A3B model and surpasses several closed-source systems including GPT-5 at 42.9%, Gemini-2.5 Pro at 44.8%, and Gemini-2.5 Flash at 41.3%.

The framework generates what the researchers describe as "high-quality cold-start tool-use trajectories" and "exploration-rich QA pairs" to train agents on complex web navigation tasks that require combining perception, tool use, and long-horizon reasoning.

Multimodal BrowseComp tasks challenge AI systems to handle compositional structure, open-world uncertainty, and multimodal integration across extended interactions with dynamic web content.

The research paper, submitted to arXiv on July 12, spans 17 pages and includes five figures demonstrating the framework's architecture and performance comparisons. The work represents an advance in training AI agents for real-world web interaction scenarios that require sophisticated multimodal reasoning capabilities.