Understanding AI Flags and the Role of Human Decision-Making

Introduction and Speaker Profiles

  • Moderator/Host: Joe.
  • Ashi Chaturvedi, PhD:
    • Program Officer for Ethics and Integrity at the American Society for Microbiology (ASM) since 20222022.
    • Handles research and publication ethics concerns.
    • Recipient of the 20242024 ISMTE Early Career Award.
    • Holds a PhD in oncological sciences from the University of Utah.
  • Alicia Hibbert:
    • Ethics Integrity Specialist at ASM since 20232023.
    • Former clinical laboratory scientist in various healthcare settings.
    • Holds a Master of Healthcare Administration from George Washington University.
  • Disclosures: Speakers have no relevant conflicts of interest (COIs). Views expressed are personal and do not necessarily represent the official positions of ASM.
  • Sponsors: River Valley Technologies, Scholastica, and Silverchair.

Fundamental Principles of AI Integration

  • Signals vs. Decisions: The central take-home message is that AI only provides signals; humans make the decisions.
  • Strengths of AI:
    • Identifying patterns.
    • Processing thousands of manuscripts rapidly.
    • Noticing details that manual review might miss.
  • Limitations of AI:
    • Lacks understanding of context.
    • Lacks understanding of intent.
    • Lacks understanding of consequences.
  • Judgment: One must not confuse AI confidence with certainty. The output is a flag, but editorial judgment remains the deciding factor.

General Framework for AI Adoption

  • Investment in Process: The primary investment is in the process, not the tool itself. The speed of AI development makes tools ephemeral, but good workflows are lasting.
  • Three-Step Workflow:
    1. Define the Problem: Identify exactly what needs to be solved (e.g., image duplication).
    2. Choose the Tool: Look for capability over features. Engage with vendors to understand strengths and limitations not listed on websites.
    3. Pilot and Refine: Test the tool in a live system and adjust based on results.
  • Defining Success: It is critical to know what success looks like and how to measure it before starting.

Case Study: Detecting Image Duplications (ImageTwin)

  • Problem Identification:
    • ASM previously had a post-acceptance figure QC process detecting splicing, joining, and manipulations.
    • They lacked a method to detect image duplications across different papers; this relied on reviewers or post-publication alerts.
  • **Tool Selection (ImageTwin): **
    • Humans cannot effectively detect duplications across large databases of published work; this is a task better suited for AI.
    • Selection criteria included cost, license requirements, workflow compatibility, ease of learning, output clarity, and vendor support.
  • Pilot Considerations:
    • Capability: Can it solve the specific problem (image duplication)?
    • Feasibility: Can it fit into the workflow without causing chaos?
    • Sustainability: Is the cost manageable for long-term use?

Workflow Refinement and Human Validation

  • Human Validation Process:
    1. AI generates a flag showing suspected duplicates.
    2. An ethics team member reviews the flags to remove false positives.
    3. Adobe Photoshop/Difference Function: To verify duplications, images are overlaid in Adobe Photoshop, and the "difference function" is applied. If the resulting image is completely black, the images are likely identical.
  • Iterative Workflow Changes:
    • Initial Pilot: Placed the tool at the post-acceptance stage.
    • Discovery: Finding issues post-acceptance led to rescinding acceptances, which harmed journal reputation and frustrated authors.
    • Refined Workflow: Moved the figure QC process to the revision/resubmission stage. This aligned with other checks like Crosscheck.
  • Scalability: The process scaled from 11 journal to 1515 journals. Scaling increased complexity, requiring more reports to review and more author communications.

Current Challenges: Paper Mills and Research Quality

  • The Problem: Journals are increasingly targeted by paper mills and low-quality research submissions.
  • Pilot Dual-Tool Strategy: To save time and gather more data, ASM is piloting two tools simultaneously.
  • Human Role in Context: Humans are needed for "nuanced thinking," such as comparing disparate data types (metaphorically comparing "pineapples to grapes").
  • Tool Comparison:
    1. Paper Mill Alarm (by Clear Skies): Analyzes network signals and connections to identify potential paper mill activity.
    2. Alchemist Review (by Alchemist): Evaluates manuscript quality and looks for integrity flags like retractions or mismatched data.
  • Metrics for Success:
    • Flags are now reviewed in minutes rather than hours.
    • Issues are typically resolved within an hour, or 33 to 55 days if author/editor contact is required.

Analysis of AI Reports

  • Paper Mill Alarm Reports: Uses a color-coded system (red, orange, green). Humans must investigate findings in the references section.
    • Caveat: A flag for a "retracted" paper might mean fraud, or it might just be a corrected author name. Humans must distinguish the severity.
    • Problematic Paper Screener: A free tool used for human confirmation of problematic papers.
  • Alchemist Review Reports: Provides a "Journal Fit Score" (11 to 1010) based on novelty, scope, and reader impact.
    • Human Judgment: What one journal considers low novelty, another might find acceptable depending on its specific standards.
  • Refinement of Objectives: ASM is investigating if the distinction between "paper mill" and "low quality" is meaningful, as low quality itself is often a marker of a paper mill.

The Human-Centered Approach and Stakeholder Management

  • Stakeholder Balancing:
    • Editors-in-Chief (EICs): Concerned about missing "diamonds in the rough."
    • IT Departments: Concerned about technical workload.
    • Editorial Teams: Require training to understand the tools and objectives.
  • Culture of Transparency: Establish boundaries for pilots. Do not feel "married" to a tool; be prepared to move on if it does not fit organizational needs.
  • Communication: Evangelize successes to get buy-in before expanding the scope of AI use.

Questions & Discussion

  • Q: What if no single tool addresses the full problem or if it introduces new challenges?
    • A: There is rarely a single answer. Focus on which trade-offs the organization is willing to manage. New tools often increase work in other areas (author communication/documentation).
  • Q: Who should make the final decision on signals?
    • A: At ASM, the ethics department handles it, but it requires back-and-forth with peer review and editorial teams. Complex problems have no easy solution.
  • Q: How do you handle non-conclusive flags?
    • A: Conduct consultations with EICs or assigned editors. Review the author's history (e.g., was a retraction 1010 years ago or recent?). Request underlying data to ease concerns. Maintain a polite, non-accusatory dialogue with authors.
  • Q: How did you check images before current tools?
    • A: Used Adobe Photoshop and Office of Research Integrity (ORI) plugins. These were effective for splicing or enhancement but not for duplication.
  • Q: Of identified issues, how many are malicious?
    • A: Most are honest errors (placeholder images, re-using controls without labeling). Out of 248248 verified duplications, only 66 (approx. 2%2\% ) resulted in revocation due to maliciousness or extreme carelessness.
  • Q: Are authors told which tools are used?
    • A: ASM identifies ImageTwin on its website. Pilot tools are not yet publicly declared until they are permanent.
  • Q: Can publishers ask for higher resolution figures at submission?
    • A: Yes, but it is hard to enforce. Image degradation often happens as authors move files between programs (PowerPoint to PDF). Higher resolutions are requested when flags occur.
  • Q: What are the next AI trends?
    • A: Tools for finding peer reviewers, managing peer review tasks, and improving document accessibility.
  • Q: Recommendations for free/affordable tools?
    • A: Office of Research Integrity (ORI) free tools, Problematic Paper Screener (for titles), Crossref (for verifying DOIs in references), and Retraction Watch (to check for previous fraud).