Exploring how interactive tools can help astronomers identify potentially incorrect labels and unusual patterns in large collections of radio astronomy data

CosmicAI Researchers Rabeya Hossain, Brian Mason, Ryan Loomis, Jeff M. Phillips, and El Kindi Rezig explored how interactive tools can help astronomers identify potentially incorrect labels and unusual patterns in large collections of radio astronomy data.

What the team did

Modern radio observatories generate large amounts of data that must be checked to ensure scientists can rely on it for their research. Automated systems can help identify unusual signals, while human experts review the data and assign labels. However, the automated results and human labels do not always agree, and experts cannot manually investigate every disagreement.

The researchers developed LabelScope, an interactive system that helps scientists explore these cases more efficiently. LabelScope brings together automated scores, human labels, and information about similar signals. It highlights cases that stand out from otherwise similar data and allows scientists to compare them with nearby examples before deciding whether they deserve further review.

The team demonstrated LabelScope using real radio astronomy signals from the Atacama Large Millimeter/submillimeter Array (ALMA).

What researchers found

The researchers showed that an unusual score or label does not always tell the full story on its own. A signal can look unusual in isolation but be completely normal when compared with other signals collected under similar conditions. Likewise, a signal may look similar to its neighbors while carrying a different label, making it worth a closer look.

LabelScope uses these surrounding signals to identify cases where a label does not match the broader pattern. It then brings those cases to an expert’s attention and shows why they were flagged, making it easier to decide which signals need further review.

Why the work matters

Radio astronomy produces thousands of signals that must be checked for problems before the data can be used confidently. Automated systems help with this process, but their scores do not always agree with the labels assigned during human review. Checking every disagreement by hand is not practical at this scale.

LabelScope helps narrow that workload by highlighting the disagreements that stand out when compared with similar signals. This gives astronomers a more focused way to review potential labeling inconsistencies and improve the quality of the data moving through the processing pipeline.

The paper was published in the Proceedings of the VLDB Endowment (PVLDB), Volume 19, and was presented as a demonstration at the 52nd International Conference on Very Large Data Bases (VLDB 2026) in Boston.

Next
Next

Clustering spectral cubes from radio astronomy