Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Jimmy Lin and colleagues at the University of Waterloo describe Project Greenhouse, an effort to build fully open and sovereign models for agentic search on modest compute, starting with Gaggle, a pointwise decoder-only reranker pre-trained from scratch.
Ask this paper
Thesis. Competitive components for agentic search can be built end to end without third-party open-weight backbones, using only commonly available datasets.
Recipe. Pre-training from scratch on the nanochat codebase, followed by supervised fine-tuning into a pointwise reranker. The released Gaggle reranker is a model soup averaging weights from four fine-tuned runs.
Compute. Pre-training used limited cycles on a single 8-GPU H100 server, and most other experiments used one or two GPUs.
Evaluation. Gaggle is compared with Qwen3-Reranker 4B and 8B and other rerankers on top-100 BM25 candidates, measured by nDCG@10 on TREC DL 2019-2023 and seven BEIR collections.
Openness. The authors formalize LLM training as a directed property hypergraph to define what fully open means, and release data, code, configurations and intermediate checkpoints for independent reproduction.
Abstract
Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.