How Researchers Navigate Accountability, Transparency, and Trust When Using AI Tools in Research

CosmicAI Researchers Sanjana Gautam, Houjiang Liu, and Matthew Lease - alongside Yujin Choi, set out to understand a question that's becoming urgent across every research discipline: when scientists lean on AI tools during the earliest, most ideation-heavy stages of their work, what happens to their sense of accountability, transparency, and trust?

To find out, the team conducted a think-aloud study with 15 researchers spanning fields from astrophysics to biochemistry, observing as they used AI-powered research tools to explore literature, synthesize findings, and generate new research ideas.

Why Early-Stage Research Is the Hardest Place to Watch

Long before a paper is written or an experiment is run, researchers are already making consequential calls like which prior studies matter, what counts as credible evidence, and which ideas are worth chasing. Those calls have always belonged to the human researchers, built up through training and experience. As AI tools increasingly assist with literature search, idea generation, and study planning, the team aimed at understanding whether those decisions are still visible and attributable, or whether they're quietly being absorbed into the AI's process. 

Watching Researchers Think Out Loud

Rather than relying on surveys or after-the-fact interviews, the team used a think-aloud protocol, asking 15 researchers to narrate their reasoning as they worked. The participants included PhD students, postdocs, a faculty member, and an industry scientist. Participants used two contrasting tools: Research Rabbit, which surfaces papers through citation-network exploration, and Elicit AI, which uses large language models to extract and synthesize findings across collections of papers. Across five tasks, participants explored unfamiliar domains, identified research threads, dug deeper into a chosen thread, and evaluated AI-generated summaries, all while thinking aloud about their strategies and discussing what they trusted, doubted, and double-checked.

Three Tensions: Accountability, Transparency, and Trust

The study's findings centered on three recurring tensions. On accountability, participants found that AI tools answer every question with the same confident tone, whether or not a clear answer actually exists. This mismatch places extra scrutiny on researchers, who remain the ones ultimately responsible for their claims. On transparency, the "black box" nature of how these tools retrieve and assemble information made it nearly impossible for researchers to establish where an answer actually came from, turning even ordinary hallucinations into transparency failures rather than simple errors. On trust, confidence in AI tools never settled into something stable; it had to be rebuilt, conversation by conversation, through constant verification. As one participant put it, describing students newer to a field: "Students would never know the applicability of information" because they never acquire, through trial and error, the skills of exploring, scoping, and validating information.

Developing workarounds to AI trust issues

Rather than walking away from these tools, participants developed workarounds to keep their own judgment in the loop. Many restricted AI use to peripheral, organizational tasks like note-taking or summarizing, while keeping core decisions about relevance and validity firmly in their own hands. Others leaned on social credibility heuristics, trusting an author's reputation as a proxy for a paper's quality, or fell back on redundant manual verification, rechecking citations and details they'd already seen the tool produce. Several described a slower, incremental process of "logical auditing," testing an AI tool's reasoning against known facts before trusting it further, and most deliberately hedged AI use to low-stakes tasks, reserving careful reading and critical evaluation for themselves.

Why It Matters

The findings suggest that AI tools, as currently designed, often add cognitive burden rather than remove it: researchers aren't simply delegating work to AI, they're spending real effort managing the uncertainty AI introduces. 

This matters most for early-career researchers, who may have the least experience to fall back on when deciding whether an AI-generated claim is sound. The authors argue that closing this gap will require more than better models. The work calls for interface and workflow design that makes provenance traceable, preserves the tools researchers already rely on (like reference managers and citation filtering), and supports the slow, situated process by which trust in a tool is actually earned.

Next
Next

Cosmic Horizons 2026: Uniting Astronomers and AI Experts at NSF NRAO