🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
200 papers · MultimodalClear filters →
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Xingang Guo, Jing Gu, Jared Lichtarge and colleagues at Scale AI (with Elorian) introduce Humanity's Sixth Sense (HSS), a benchmark for the intuitive visual reasoning people perform at a glance.

01Multimodal
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

02Memory
Pistis Technical Report

Pistis Technical Report

The Pistis team at ByteDance introduces 27B and 9B multimodal models on Qwen3.6 and Qwen3.5, trained with interleaved on-policy distillation and RL, plus an automatic harness optimizer.

03Multimodal
Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen3.8-Omni: Towards Native Omni-Modal Agents

The Qwen Team at Alibaba releases Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agentic work across text, audio and video, together with two open-source frameworks for running it as an agent.

04Agents
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Zhangxuan Gu and colleagues present LLaDA-UI, a 16.7B-parameter mixture-of-experts GUI agent built on a block-wise diffusion language backbone, and test whether diffusion decoding can support capable multimodal GUI agents.

05Agents
A frontend-backend architecture for tool calls in full-duplex speech models

A frontend-backend architecture for tool calls in full-duplex speech models

Ke Hu and colleagues at NVIDIA give a full-duplex speech-to-speech model tool-calling ability by having the speech frontend emit a delegation token and hand streaming transcripts to a text backend LLM, rather than teaching the duplex model to call tools itself.

06Multimodal
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI releases DeepSeek-V4.1-Flash, a 552B-parameter multimodal MoE built around a Causal Encoder-Decoder architecture that activates 16B parameters per decode token and 8B per prefill token, and cuts the resident KV cache to 890 bytes per token.

07Memory
NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

Johnny Greco and colleagues at NVIDIA describe NeMo Data Designer, an open-source framework for multimodal synthetic data generation in which humans or agents declare each dataset column in an inspectable configuration.

08Data
RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

Yunxiang Zhang, Yan Chen and colleagues (Beihang University) present RepoAtlas, a training-free module that gives coding agents an evolving visual and textual view of the relevant part of a repository code graph.

09Code
Online Video Agent Harness for Long Video Understanding

Online Video Agent Harness for Long Video Understanding

Sen Yang and colleagues at Baidu build VideoXAgent, an online agent harness for long videos that plans from the query, calls expert tools on demand and aggregates evidence, instead of packing dense frames into one context.

10Agents
MindTopo: Can Foundation Models Reason in Topological Space?

MindTopo: Can Foundation Models Reason in Topological Space?

Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

11Reasoning
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Jiaqiang Li, Tao Gui and colleagues (Fudan NLP Group with Shanghai AI Laboratory) build Sci-MMR, a benchmark that checks whether multimodal research agents recover the full evidence chain behind a scientific answer, and find answer accuracy runs more than 20 points ahead of evidence recovery.

12Multimodal
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Minghao Guo, Meng Cao and colleagues (MBZUAI and USTC) build Mr.LHDR, a deep research benchmark whose questions require long chains of dependent, multimodal evidence, and find that the best system fully completes only about a third of them.

13Multimodal
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Minji Kim, Jihyoung Jang and Hyounghun Kim at POSTECH argue that non-compliance in vision-language models is evaluated at the wrong granularity, and build a benchmark where a single query mixes answerable content with content that should be withheld.

14Multimodal
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Wonje Jeung and colleagues at Yonsei University, with Carnegie Mellon, show that vision-language models used as reward functions for robot learning give different rewards to the same trajectory when the goal instruction is paraphrased, and release a benchmark that measures it.

15Multimodal
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Sanyuan Chen and colleagues at FAIR at Meta present Text-AB, a 3B latent-diffusion speech model that removes the forced-alignment stage from the Audiobox line and handles dubbing and two-speaker dialogue in one system.

16Safety
MedQA-MM: Shortcuts Behind Medical Visual Reasoning

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Benlu Wang and colleagues at UMass Amherst and Yale separate the answer from the route that produced it in medical multimodal MCQs, and find that scores substantially overstate image reasoning.

17Evaluation
Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Xingming Long and colleagues introduce NTEP, an annotation scheme that names the necessary external evidence and the tool calls that must produce it, and NTEP-R, a reward that pays the agent per tool call rather than only on the final answer.

18Agents
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Zhaoyuan Huang and colleagues at Shanghai Jiao Tong and Ant Group ask whether GUI agents know when not to act, build CONFLICTGUI to measure it, and find severe execution-biased overcompliance across five widely used agents.

19Agents
DataSpace

DataSpace

Real organizational analytics scatters evidence across databases, structured files, long documents, and video. Existing benchmarks isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic scoring untested together.

20Evaluation
GAMUT

GAMUT

Most factuality evaluation measures precision, whether the claims in an answer are correct. This Meta AI work targets the harder and mostly ignored half, completeness, meaning whether an answer covers everything it should, and packages it as the GAMUT benchmark.

21Evaluation
RoboTTT

RoboTTT

Recent robot foundation models run on single-step or short-history context, a strange way to attempt a five-minute assembly task. RoboTTT, from NVIDIA with Stanford and UT Austin, integrates test-time training into vision-language-action policies to scale visuomotor context to 8K timesteps, three orders of magnitude past prior policies, without growing inference latency. The longer context unlocks one-shot in-context imitation from human video, on-the-fly policy improvement, and robustness to perturbations. It improves overall performance by 87% over a single-step baseline, fully completes a ten-stage assembly task that no baseline finishes, and gains 62% from pretraining with 8K rather than 1K timesteps.

22Robotics
LingBot-World 2.0

LingBot-World 2.0

Most world models fall apart after a few seconds, smearing textures and warping geometry as errors compound frame to frame. LingBot-World 2.0 from Robbyant holds 720p at 60 fps for a full hour of interaction and ships fully open.

23Multimodal
LingBot-VLA 2.0

LingBot-VLA 2.0

LingBot-VLA 2.0 is an open-source generalist embodied model from Robbyant, trained across 20 robot configurations from single-arm rigs to humanoids like Unitree G1 and Fourier GR-2. It packs 60,000 hours of curated data, 50,000 hours of real-robot trajectories plus 10,000 hours of egocentric human video, into one policy that also predicts future depth and semantic features before it acts. On 9 GM-100 tabletop tasks it beats π0.5 and GR00T N1.7 across two robot platforms and stays ahead on long-horizon mobile tasks, running at about 130 ms on a single RTX 4090D with open-sourced post-training code.

24Robotics
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026