🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papersIssue 107 of 176

The week of Apr 14 – Apr 20, 2025

10 papers, hand-picked and summarised.

GUI-R1

GUI-R1

Researchers from the National University of Singapore and the Chinese Academy of Sciences introduce GUI-R1, a reinforcement learning (RL) framework aimed at improving graphical user interface (GUI) agents through unified action-space modeling. Key insights include:

01Reinforcement Learning
Scaling Reasoning in Diffusion LLMs via RL

Scaling Reasoning in Diffusion LLMs via RL

Proposes d1, a two‑stage recipe that equips masked diffusion LLMs with strong step‑by‑step reasoning.

02Reasoning
Enhancing Non-Reasoning Models with Reasoning Models

Enhancing Non-Reasoning Models with Reasoning Models

Researchers explore how to distill reasoning-intensive outputs (answers and explanations) from top-tier LLMs into more lightweight models that don’t explicitly reason step by step. By fine-tuning smaller models on the high-quality final answers (and optionally summarized thinking traces) from advanced reasoning models, they demonstrate consistent performance boosts across multiple benchmarks.

03Reasoning
AgentA/B

AgentA/B

AgentA/B is a fully automated A/B testing framework that replaces live human traffic with large-scale LLM-based agents. These agents simulate realistic, intention-driven user behaviors on actual web environments, enabling faster, cheaper, and risk-free UX evaluations — even on real websites like Amazon. Key Insights:

04Agents
Reasoning Models Can Be Effective Without Thinking

Reasoning Models Can Be Effective Without Thinking

This paper challenges the necessity of long chain-of-thought (CoT) reasoning in LLMs by introducing a simple prompting method called NoThinking, which bypasses explicit "thinking" steps. Surprisingly, NoThinking performs comparably to or better than traditional reasoning under comparable or even lower compute budgets, especially when paired with parallel decoding and best-of-N selection. Key Insights:

05Reasoning
SocioVerse

SocioVerse

Researchers from Fudan University and collaborators propose SocioVerse, a large-scale world model for social simulation using LLM agents aligned with real-world user behavior. Key ideas include:

06Safety
DocAgent

DocAgent

Researchers from Meta AI present DocAgent, a tool‑integrated, dependency‑aware framework that turns large, complex codebases into well‑written docstrings. Key ideas include:

07Agents
SWE-PolyBench

SWE-PolyBench

SWE-PolyBench is a new multi-language benchmark for evaluating coding agents on real-world software tasks across Java, JavaScript, TypeScript, and Python. It introduces execution-based assessments, syntax tree metrics, and reveals that current agents struggle with complex tasks and show inconsistent performance across languages.

08Evaluation
A Survey of Frontiers in LLM Reasoning

A Survey of Frontiers in LLM Reasoning

This survey categorizes LLM reasoning methods by when reasoning occurs (inference-time vs. training) and the system's architecture (standalone vs. agentic or multi-agent). It highlights trends like learning-to-reason (e.g., DeepSeek-R1) and agentic workflows (e.g., OpenAI Deep Research), covering prompt engineering, output refinement, and learning strategies such as PPO and verifier training.

09Agents
Advances in Embodied Agents, Smart Cities, and Earth Science

Advances in Embodied Agents, Smart Cities, and Earth Science

This paper surveys how spatial intelligence manifests across disciplines—from embodied agents to urban and global systems—by connecting human spatial cognition with how LLMs handle spatial memory, representations, and reasoning. It offers a unifying framework to bridge research in AI, robotics, urban planning, and earth science, highlighting LLMs’ evolving spatial capabilities and their interdisciplinary potential.

10Robotics
Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack