AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
Xingang Guo, Jing Gu, Jared Lichtarge and colleagues at Scale AI (with Elorian) introduce Humanity's Sixth Sense (HSS), a benchmark for the intuitive visual reasoning people perform at a glance.

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

Pistis Technical Report
The Pistis team at ByteDance introduces 27B and 9B multimodal models on Qwen3.6 and Qwen3.5, trained with interleaved on-policy distillation and RL, plus an automatic harness optimizer.

Qwen3.8-Omni: Towards Native Omni-Modal Agents
The Qwen Team at Alibaba releases Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agentic work across text, audio and video, together with two open-source frameworks for running it as an agent.

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Zhangxuan Gu and colleagues present LLaDA-UI, a 16.7B-parameter mixture-of-experts GUI agent built on a block-wise diffusion language backbone, and test whether diffusion decoding can support capable multimodal GUI agents.

A frontend-backend architecture for tool calls in full-duplex speech models
Ke Hu and colleagues at NVIDIA give a full-duplex speech-to-speech model tool-calling ability by having the speech frontend emit a delegation token and hand streaming transcripts to a text backend LLM, rather than teaching the duplex model to call tools itself.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-AI releases DeepSeek-V4.1-Flash, a 552B-parameter multimodal MoE built around a Causal Encoder-Decoder architecture that activates 16B parameters per decode token and 8B per prefill token, and cuts the resident KV cache to 890 bytes per token.

NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation
Johnny Greco and colleagues at NVIDIA describe NeMo Data Designer, an open-source framework for multimodal synthetic data generation in which humans or agents declare each dataset column in an inspectable configuration.

RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views
Yunxiang Zhang, Yan Chen and colleagues (Beihang University) present RepoAtlas, a training-free module that gives coding agents an evolving visual and textual view of the relevant part of a repository code graph.

Online Video Agent Harness for Long Video Understanding
Sen Yang and colleagues at Baidu build VideoXAgent, an online agent harness for long videos that plans from the query, calls expert tools on demand and aggregates evidence, instead of packing dense frames into one context.

MindTopo: Can Foundation Models Reason in Topological Space?
Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Jiaqiang Li, Tao Gui and colleagues (Fudan NLP Group with Shanghai AI Laboratory) build Sci-MMR, a benchmark that checks whether multimodal research agents recover the full evidence chain behind a scientific answer, and find answer accuracy runs more than 20 points ahead of evidence recovery.

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Minghao Guo, Meng Cao and colleagues (MBZUAI and USTC) build Mr.LHDR, a deep research benchmark whose questions require long chains of dependent, multimodal evidence, and find that the best system fully completes only about a third of them.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Minji Kim, Jihyoung Jang and Hyounghun Kim at POSTECH argue that non-compliance in vision-language models is evaluated at the wrong granularity, and build a benchmark where a single query mixes answerable content with content that should be withheld.

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
Wonje Jeung and colleagues at Yonsei University, with Carnegie Mellon, show that vision-language models used as reward functions for robot learning give different rewards to the same trajectory when the goal instruction is paraphrased, and release a benchmark that measures it.

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
Sanyuan Chen and colleagues at FAIR at Meta present Text-AB, a 3B latent-diffusion speech model that removes the forced-alignment stage from the Audiobox line and handles dubbing and two-speaker dialogue in one system.

MedQA-MM: Shortcuts Behind Medical Visual Reasoning
Benlu Wang and colleagues at UMass Amherst and Yale separate the answer from the route that produced it in medical multimodal MCQs, and find that scores substantially overstate image reasoning.

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
Xingming Long and colleagues introduce NTEP, an annotation scheme that names the necessary external evidence and the tool calls that must produce it, and NTEP-R, a reward that pays the agent per tool call rather than only on the final answer.

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
Zhaoyuan Huang and colleagues at Shanghai Jiao Tong and Ant Group ask whether GUI agents know when not to act, build CONFLICTGUI to measure it, and find severe execution-biased overcompliance across five widely used agents.

DataSpace
Real organizational analytics scatters evidence across databases, structured files, long documents, and video. Existing benchmarks isolate structured querying, retrieval, or open-ended analysis, leaving heterogeneous evidence discovery, complete tabular outputs, and deterministic scoring untested together.

GAMUT
Most factuality evaluation measures precision, whether the claims in an answer are correct. This Meta AI work targets the harder and mostly ignored half, completeness, meaning whether an answer covers everything it should, and packages it as the GAMUT benchmark.

RoboTTT
Recent robot foundation models run on single-step or short-history context, a strange way to attempt a five-minute assembly task. RoboTTT, from NVIDIA with Stanford and UT Austin, integrates test-time training into vision-language-action policies to scale visuomotor context to 8K timesteps, three orders of magnitude past prior policies, without growing inference latency. The longer context unlocks one-shot in-context imitation from human video, on-the-fly policy improvement, and robustness to perturbations. It improves overall performance by 87% over a single-step baseline, fully completes a ten-stage assembly task that no baseline finishes, and gains 62% from pretraining with 8K rather than 1K timesteps.

LingBot-World 2.0
Most world models fall apart after a few seconds, smearing textures and warping geometry as errors compound frame to frame. LingBot-World 2.0 from Robbyant holds 720p at 60 fps for a full hour of interaction and ships fully open.

LingBot-VLA 2.0
LingBot-VLA 2.0 is an open-source generalist embodied model from Robbyant, trained across 20 robot configurations from single-arm rigs to humanoids like Unitree G1 and Fourier GR-2. It packs 60,000 hours of curated data, 50,000 hours of real-robot trajectories plus 10,000 hours of egocentric human video, into one policy that also predicts future depth and semantic features before it acts. On 9 GM-100 tabletop tasks it beats π0.5 and GR00T N1.7 across two robot platforms and stays ahead on long-horizon mobile tasks, running at about 130 ms on a single RTX 4090D with open-sourced post-training code.