Making astronomical archives more accessible through AI: STABLE (Summer Team for Astronomical Benchmarking & LLM Engineering) highlights

Three CosmicAI research projects are being highlighted this month that are unified by the CosmicAI seed funding project called STABLE (Summer Team for Astronomical Benchmarking & LLM Engineering), which is part of an ongoing effort to make astronomical archives more accessible through AI.

The STABLE program has three pillars: LLM benchmarking, infrastructure enhancement, and AI agent development.  A cohort of three interns joined a team of nine CosmicAI researchers over a period of about three months at NSF NRAO and NSF NOIRLab.  Through the work in parallel, the team has explored deployment of their tools at TACC and NOIRLab Astro Data Lab, presented a paper at the Conference for Physics and AI, and prepared three posters for the CosmicAI Cosmic Horizons conference.  Deliverables from the team include Github repositories, documentation of tools, and benchmarking standards, as well as ongoing conversation with the CosmicAI Explorable Universe group about lessons learned and next steps.

MANNA – An MCP Server for Accessing Astronomical Archives

Watch the short presentation by Dan Gause, postbac researcher at NSF NRAO, as he describes his research in building an AI-agent interface to the world's astronomical data archives. The CosmicAI team included Brian Mason, Adele Plunkett, Stephanie Juneau, and Robert Nikutta.

What they did

The team built MANNA (MCP Architecture for NOIRLab and NSF NRAO Archives), an open-source server that lets an LLM assistant query professional astronomy archives directly. It exposes the IVOA standard protocols – TAP/ADQL, image search, cone search, registry discovery, name resolution – across NOIRLab Astro Data Lab and NSF NRAO/ALMA, encoding each archive's real-world quirks so the model queries them correctly. 

Why it matters

The significance of the work lies in the gap between "the data is public" and "the data is usable." Petabytes of astronomical data sit behind open, standardized interfaces that still require ADQL and archive-specific expertise to reach. MANNA moves that expertise into the tooling, so a plain-language question returns real observations instead of a model's training data — and any IVOA-compliant archive can be added as a single file.


Quasar: An AI Powered Research Assistant for Astronomy Archive Workflows

Watch the short presentation by Adam Zacharia Anil, Research Intern at CosmicAI, as he describes his research in building Quasar, an open source AI research assistant for natural language access to Astronomy Archives. The CosmicAI team included Adam Zacharia Anil, Adele L. Plunkett, Brian S. Mason, and Stella S. R. Offner.

What they did

The team developed Quasar as a domain specialized research assistant that converts natural language questions into structured archive searches, retrieves relevant observations and scientific literature, supports analysis, and creates reproducible research outputs.

Why it matters

The significance of the work lies in making large Astronomy Archives easier to use and reducing the specialized technical knowledge needed to move from a scientific question to relevant data and analysis while keeping the workflow transparent and reproducible.

Benchmarking LLM Reliability in Multiwavelength Astronomy

Zoie Telkamp, a PhD candidate at the University of Virginia, describes the project as a framework for testing whether large language models like Claude can be trusted to help astronomers access and use data across different archives and wavelength regimes.

The CosmicAI team included Adele Plunkett, Brian Mason, Stephanie Juneau, Robert Nikutta, Stella Offner, Dan Gause, Adam Anil, Ryan Loomis, Eric Murphy, Greg Durrett, and Paul Torrey.

What they did

The team developed a reusable benchmarking framework that evaluates how well LLMs perform on two kinds of tasks: multiwavelength data discovery (finding the right data across archives) and cross-regime domain expertise (answering conceptual questions that come up when working in an unfamiliar wavelength regime). Data discovery tasks are executed and evaluated using the Harbor task framework with pytest verification against archive query results, and domain expertise tasks are evaluated against expert-written reference answers using a structured rubric.

Why it matters

The significance of the work lies in giving astronomers a repeatable way to robustly assess LLM performance on multiwavelength astronomy tasks, pinpointing areas of strength and identifying gaps in archive and domain knowledge. The framework and prompt set are being released publicly so other researchers can extend the work to additional prompts and run their own evaluations as models continue to evolve.

Next
Next

“Vector Summary Pseudo Posteriors for Simulation-Based Inference with Applications to Cosmology” presented at the STAI-X conference at Harvard