Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

Haiyang Xu, Xi Zhang, Haowei Liu, Junyang Wang, Zhaozai Zhu, Shengjie Zhou, Xuhao Hu, Feiyu Gao, Junjie Cao, Zihua Wang, Zhiyuan Chen, Jitong Liao, Qi Zheng, Jiahui Zeng, Ze Xu, Shuai Bai, Junyang Lin, Jingren Zhou, Ming Yan · Feb 15, 2026 · Citations: 0

General Long Horizon Multi Agent Simulation Env Web Browsing

Open arXiv RSS feed

Abstract

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge collaboration and real-time interaction. GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena; (2) on grounding tasks, it obtains 80.3 on ScreenSpotPro; (3) on tool-calling tasks, it obtains 47.6 on OSWorld-MCP, and 46.8 on MobileWorld; (4) on memory and knowledge tasks, it obtains 75.5 on GUI-Knowledge Bench. GUI-Owl-1.5 incorporates several key innovations: (1) Hybird Data Flywheel: we construct the data pipeline for UI understanding and trajectory generation based on a combination of simulated environments and cloud-based sandbox environments, in order to improve the efficiency and quality of data collection. (2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and multi-agent adaptation; (3) Multi-platform Environment RL Scaling: We propose a new environment RL algorithm, MRPO, to address the challenges of multi-platform conflicts and the low training efficiency of long-horizon tasks. The GUI-Owl-1.5 models are open-sourced, and an online cloud-sandbox demo is available at https://github.com/X-PLUG/MobileAgent.

HFEPX Relevance Assessment

This paper has direct human-feedback and/or evaluation protocol signal and is likely useful for eval pipeline design.

Eval-Fit Score

27/100 • Low

Treat as adjacent context, not a core eval-method reference.

Human Feedback Signal

Not explicit in abstract metadata

Evaluation Signal

Detected

HFEPX Fit

High-confidence candidate

If you are doing eval pipeline work, start here:

Human Eval Hub LLM-as-Judge Hub Pairwise Preference Hub Tool-Use Eval Hub

Human Data Lens

Uses human feedback: No
Feedback types: None
Rater population: Unknown
Unit of annotation: Trajectory
Expertise required: General
Extraction source: Persisted extraction

Evaluation Lens

Evaluation modes: Simulation Env
Agentic eval: Long Horizon, Multi Agent, Web Browsing
Quality controls: Not reported
Confidence: 0.50
Flags: ambiguous

Protocol And Measurement Signals

Benchmarks / Datasets

WebArenaOSWorld

Reported Metrics

No metric terms were extracted from the available abstract.

Research Brief

Deterministic synthesis

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge… HFEPX signals include Simulation Env, Long Horizon, Multi Agent with confidence 0.50. Updated from current HFEPX corpus.

Generated Mar 3, 2026, 6:45 PM · Grounded in abstract + metadata only

Key Takeaways

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms…
GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld,…

Researcher Actions

Treat this as method context, then pivot to protocol-specific HFEPX hubs.
Cross-check benchmark overlap: WebArena, OSWorld.
Verify metric definitions before comparing against your eval pipeline.

Caveats

Generated from title, abstract, and extracted metadata only; full-paper implementation details are not parsed.
Extraction confidence is probabilistic and should be validated for critical decisions.

Recommended Queries

human-eval protocol design agent eval benchmark comparison inter-rater agreement adjudication

Research Summary

Contribution Summary

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge…
GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena; (2) on grounding tasks, it obtains…
(2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and…

Why It Matters For Eval

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge…
(2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and…

Researcher Checklist

Gap: Human feedback protocol is explicit

No explicit human feedback protocol detected.
Pass: Evaluation mode is explicit

Detected: Simulation Env
Gap: Quality control reporting appears

No calibration/adjudication/IAA control explicitly detected.
Pass: Benchmark or dataset anchors are present

Detected: WebArena, OSWorld
Gap: Metric reporting is present

No metric terms extracted.

Related Papers

Papers are ranked by protocol overlap, extraction signal alignment, and semantic proximity.

Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning Protocol Overlap

Citations: 0 Relevance: 11.20 Shared tag: Simulation EnvShared tag: Long HorizonShared tag: Web Browsing
- Shared 3 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
Efficient Hierarchical Any-Angle Path Planning on Multi-Resolution 3D Grids Protocol Overlap

Citations: 0 Relevance: 11.20 Shared tag: Simulation EnvShared tag: Long HorizonShared tag: Web Browsing
- Shared 3 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation Protocol Overlap

Citations: 0 Relevance: 11.20 Shared tag: Simulation EnvShared tag: Long HorizonShared tag: Web Browsing
- Shared 3 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
CoAct-1: Computer-using Multi-Agent System with Coding Actions Protocol Overlap

Citations: 0 Relevance: 8.60 Shared tag: Long HorizonShared tag: Multi Agent
- Shared 2 HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared benchmark mentions
"Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Web Browsing
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
A Survey on the Optimization of Large Language Model-based Agents Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Long Horizon
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Long Horizon
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
Architecting AgentOS: From Token-Level Context to Emergent System-Level Intelligence Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Multi Agent
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Long Horizon
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Long Horizon
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
BrowseComp-$V^3$: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Web Browsing
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup
Contextual Safety Reasoning and Grounding for Open-World Robots Protocol Overlap

Citations: 0 Relevance: 7.50 Shared tag: Simulation EnvShared tag: Web Browsing
- Shared 2 HFEPX protocol tags
- Aligned evaluation mode
- Aligned agent-evaluation setup

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote