
LLaSM (Large Language and Speech Model)
A combined language-and-speech model trained with cross-modal conversational abilities.

SAM-Med2D
Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

Vector Search with OpenAI Embeddings
Argues, via empirical analysis, that dedicated vector databases aren't necessarily required for modern AI-stack search applications.

Graph of Thoughts (GoT)
Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

MVDream
ByteDance's MVDream is a multi-view diffusion model that generates geometrically consistent images from multiple viewpoints given a text prompt.

Nougat
Meta's Nougat is a visual transformer for "Neural Optical Understanding for Academic documents" that converts PDFs to LaTeX/Markdown.

FacTool
A tool-augmented framework for detecting factual errors in LLM-generated text.

AnomalyGPT
Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

FaceChain
Alibaba's FaceChain is a personalized portrait generation framework that produces identity-preserving portraits from just a handful of input photos.

Qwen-VL
Alibaba's Qwen-VL is a large-scale vision-language model family with strong performance across captioning, VQA, and visual localization.