GUI navigation · ECCV 2026

AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation

Shengjie Jin, Zelong Sun, Hengbo Xu, Yanbiao Ma, Zhiwu Lu

Gaoling School of Artificial Intelligence, Renmin University of China

We identify informational asymmetry in GUI navigation and introduce a unified framework driven by the Cognitive State Anchor. Its asymmetric memory unifies intra-trial correction and cross-trial experience distillation by preserving the visual evidence that explains failures.

Existing GUI agents compress away unexpected visual states, while AnchorGUI retains mismatch evidence through Cognitive State Anchors for correction and learning.
Figure 1. Compressing visual histories can erase evidence of unexpected states. AnchorGUI keeps screenshots for mismatched transitions and uses them for correction within a trial and experience distillation across trials.Enlarge ↗

Which GUI states need visual evidence?

GUI agents need to correct their decisions within a trial and learn from failures across trials. Both depend on identifying the visual evidence that explains an unexpected outcome.

Consider an agent searching for a video that instead lands on a “Welcome to Chrome” setup screen. If history compression removes this visual evidence, the agent can repeatedly press Back and relaunch the app without recognizing the loop. Retaining every screenshot inflates context and buries failure-causing actions among visually similar steps.

The key observation is informational asymmetry. A transition that matches the agent’s expectation can often be summarized in text. A mismatch benefits from screenshots that preserve the causal evidence needed for diagnosis. This same distinction helps the agent correct its next action and extract useful experience after a failed trial.

Cognitive State Anchors and asymmetric memory

AnchorGUI policy, execution, evaluator, and Cognitive State Anchor drive asymmetric memory updates within trials and experience distillation across failed trials.
Cognitive State Anchors drive intra-trial correction (top) and cross-trial experience distillation (bottom). Source: paper, Figure 2.Enlarge ↗

The Cognitive State Anchor

The policy predicts an action and an expected GUI transition. After execution, the evaluator compares the expected transition with the observed change. The Cognitive State Anchor (CSA) records the expected transition, observed transition, and a binary match or mismatch verdict.

One memory, two learning timescales

Within a trial, a sliding window retains screenshots for mismatched transitions and lightweight text records for matched ones. The agent receives visually grounded evidence to correct its next decision.

Across trials, a distiller uses the asymmetric memory of a failed trial to extract reusable experience. Preserved mismatch evidence focuses credit assignment on likely failure steps, and the resulting experience guides the next attempt through in-context adaptation, without parameter updates.

GUI navigation and cross-trial learning

On AndroidWorld with Qwen3-VL-8B, AnchorGUI achieves 57.3% single-attempt success, compared with 43.7% for ReAct using the same backbone. The complete tables below cover offline action prediction, online task execution, and cross-trial learning. Table 3 also includes a controlled comparison with the AnchorGUI intra-trial module fixed.

Table 1 · Offline GUI navigation benchmarks

Scroll horizontally to view all columns →

MethodModelAITZAndroid-ControlGUI-Odyssey
TM (%)EM (%)TM (%)EM (%)TM (%)EM (%)
Fine-tuning Methods
OS-GenesisOS-Genesis-7B20.08.4565.944.411.73.63
AguvisAguvis-7B35.719.065.654.226.713.5
OdysseyAgentOdysseyAgent59.231.658.832.790.873.7
UI-TARSUI-TARS-7B71.755.368.560.878.857.3
Zero-shot Methods
GPT-4oGPT-4o70.035.363.130.937.514.2
ReActQwen3-VL-8B71.457.571.870.075.458.9
ReSumQwen3-VL-8B69.951.972.370.576.554.3
TruncationQwen3-VL-8B71.056.772.070.578.360.3
Self-ReflectionQwen3-VL-8B73.958.071.970.380.160.7
AnchorGUI (Ours)Qwen3-VL-8B75.259.174.673.080.660.6

Single-trial evaluation. TM: Type Match of the action category. EM: Exact Match of the category and all action parameters. Bold values follow the paper’s best results within each method category.

Table 2 · Single-trial performance on AndroidWorld

Scroll horizontally to view all columns →

MethodModelSR (%)
Fine-tuning Methods
GUI-CriticGPT-4o + GUI-Critic-R129.4
UGroundGPT-4o + UGround32.8
AguvisGPT-4o + Aguvis-7B37.1
Aria-UIGPT-4o + Aria-UI44.8
UI-TARSUI-TARS-72B-SFT46.6
Agent-S2Claude-3.7-Sonnet + UI-TARS-72B-DPO54.3
Zero-shot Methods
GPT-4oGPT-4o34.5
AndroidGenGPT-4o46.8
MobileUseQwen2.5-VL-7B21.6
MobileUseQwen2.5-VL-32B44.4
ReActQwen3-VL-8B43.7
ReSumQwen3-VL-8B45.6
TruncationQwen3-VL-8B44.7
Self-ReflectionQwen3-VL-8B47.7
M3AQwen3-VL-8B39.8
Mobile-Agent-v3Qwen3-VL-8B55.2
Chain-of-MemoryQwen3-VL-8B46.8
AnchorGUI (Ours)Qwen3-VL-8B57.3

SR: task Success Rate. Models are retained for every method, including results using different backbones. Bold values follow the paper’s best results within each category.

Table 3 · Cross-trial learning and token efficiency on AndroidWorld

Scroll horizontally to view all columns →

MethodSuccess rate (%)Average tokens
SR@1SR@2SR@3Δ (T3−T1)Intra / stepCross / trial
Baseline intra-trial + various cross-trial strategies
ReAct + Episodic Memory43.749.153.4+9.7491.4k-
ReAct + Reflexion43.753.055.2+11.532.7k59.2k
Self-Reflection + Episodic Memory47.751.654.2+6.5098.5k-
Self-Reflection + Reflexion47.752.556.8+9.1338.1k60.7k
AnchorGUI intra-trial + various cross-trial strategies
Memoryless Retry57.361.262.9+5.5113.5k-
Episodic Memory57.360.562.6+5.2978.9k-
Reflexion57.363.164.4+7.0213.6k66.5k
Asymmetric Distillation (Ours)57.365.469.2+11.913.7k24.4k

SR@k is cumulative success within at most k attempts. Δ values are reported as in the paper, based on unrounded results. Intra tokens include policy generation and step-level verification; Cross tokens measure distillation per trial. The lower block fixes the AnchorGUI intra-trial module. Bold values follow the paper.

Experimental setting

Evaluation covers AITZ, Android-Control, GUI-Odyssey, and AndroidWorld. The online comparison uses Qwen3-VL-8B and at most three attempts. Cross-trial adaptation extracts in-context experience without parameter updates; token costs are reported separately in the paper.

BibTeX

@inproceedings{jin2026anchorgui,
  title = {{AnchorGUI}: Asymmetric Memory for Dual-Scale Learning in {GUI} Navigation},
  author = {Shengjie Jin and Zelong Sun and Hengbo Xu and Yanbiao Ma and Zhiwu Lu},
  booktitle = {Computer Vision -- ECCV 2026},
  year = {2026},
  series = {Lecture Notes in Computer Science},
  volume = {17038},
  pages = {97--113},
  publisher = {Springer Nature Switzerland},
  address = {Cham},
  doi = {10.1007/978-3-032-37447-9_6},
  url = {https://doi.org/10.1007/978-3-032-37447-9_6}
}

Research figure

Open original ↗Full figure