
LLaSM (Large Language and Speech Model)
A combined language-and-speech model trained with cross-modal conversational abilities.

SAM-Med2D
Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

Vector Search with OpenAI Embeddings
Argues, via empirical analysis, that dedicated vector databases aren't necessarily required for modern AI-stack search applications.

Graph of Thoughts (GoT)
Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

MVDream
ByteDance's MVDream is a multi-view diffusion model that generates geometrically consistent images from multiple viewpoints given a text prompt.

Nougat
Meta's Nougat is a visual transformer for "Neural Optical Understanding for Academic documents" that converts PDFs to LaTeX/Markdown.

FacTool
A tool-augmented framework for detecting factual errors in LLM-generated text.

AnomalyGPT
Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

FaceChain
Alibaba's FaceChain is a personalized portrait generation framework that produces identity-preserving portraits from just a handful of input photos.

Qwen-VL
Alibaba's Qwen-VL is a large-scale vision-language model family with strong performance across captioning, VQA, and visual localization.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack