Applying machine-learning methods for scalable and validated radio astronomy data preparation and quality assurance

CosmicAI Researchers Dr. Ci Xue, Dr. Brian Mason, Gazi Rakib, Prof. Jeff Phillips, Dr. Ryan Loomis, Dr. Omkar Bait, Dr. Eric Murphy explored applying machine-learning methods for scalable and validated radio astronomy data preparation and quality assurance. The team developed and implemented a supervised classifier framework for identifying calibration anomalies in interferometric radio observations. 

The team collaborated with Tristan Ashton, Ilsang Yoon, John Hibbard, and Ignacio Toledo.

What is radio astronomy data calibration

Radio telescopes do not record the sky signal perfectly. Instrumental and atmospheric effects can change the measured signal, so the observations must be calibrated before they can be used for scientific analysis. Calibration measures these effects and applies corrections to the data. One important step is bandpass calibration, which corrects the instrumental response as a function of observing frequency. If corrupted calibration solutions are not identified and flagged, they can propagate through downstream processing and introduce artifacts into science products.

What the team did 

The team used  an XGBoost model that sits on top of multiple features and specialized scan statistic models to identify anomalies in amplitude solutions from ALMA bandpass calibrations. This framework is general with respect to feature construction and anomaly types, with the current implementation focusing on the amplitude portions of complex bandpass solutions. Using TACC computing resources, the team assembled labeled training data from selected ALMA Cycle 9 observations containing more than 84,000 bandpass calibration solutions. The primary features are based on an adapted interval anomaly detection method, Scan Statistics, developed by Gazi Rakib and Jeff Phillips. Leveraging complementary domain knowledge, these statistical measurements support the XGBoost classifier to distinguish actionable instrumental artifacts from expected atmospheric effects.

What researchers found

By treating the identification of calibration anomalies as a supervised classification task,   the machine-learning classifier learns and optimizes complex, non-linear decision boundaries across multiple features. When evaluated against expert-validated labels, both false negatives, corresponding to missed anomalous solutions, and false positives, corresponding to false alarms, accounted for less than 0.1% of the total sample set. In a benchmark against the existing observatory pipeline heuristic, this machine-learning classifier reduced both false-positive and false-negative rates by more than 70%.

Why the work matters

In observatory data processing pipelines, identifying calibration corruptions and anomalies often relies on human inspection or expert-crafted heuristics with manually tuned thresholds. These approaches become progressively difficult to scale as observational data volumes grow.

By improving automated flagging decisions and calibration quality assurance, this work provides a scalable anomaly detection framework that can help observatories deliver science-ready data products more efficiently and be adapted for next-generation radio facilities

Next
Next

How the First Galaxies Ionized the Neutral Hydrogen at the Cosmic Scales