Understanding AI Flags and the Role of Human Decision-Making
Introduction and Speaker Profiles
- Moderator/Host: Joe.
- Ashi Chaturvedi, PhD:
- Program Officer for Ethics and Integrity at the American Society for Microbiology (ASM) since 2022.
- Handles research and publication ethics concerns.
- Recipient of the 2024 ISMTE Early Career Award.
- Holds a PhD in oncological sciences from the University of Utah.
- Alicia Hibbert:
- Ethics Integrity Specialist at ASM since 2023.
- Former clinical laboratory scientist in various healthcare settings.
- Holds a Master of Healthcare Administration from George Washington University.
- Disclosures: Speakers have no relevant conflicts of interest (COIs). Views expressed are personal and do not necessarily represent the official positions of ASM.
- Sponsors: River Valley Technologies, Scholastica, and Silverchair.
Fundamental Principles of AI Integration
- Signals vs. Decisions: The central take-home message is that AI only provides signals; humans make the decisions.
- Strengths of AI:
- Identifying patterns.
- Processing thousands of manuscripts rapidly.
- Noticing details that manual review might miss.
- Limitations of AI:
- Lacks understanding of context.
- Lacks understanding of intent.
- Lacks understanding of consequences.
- Judgment: One must not confuse AI confidence with certainty. The output is a flag, but editorial judgment remains the deciding factor.
General Framework for AI Adoption
- Investment in Process: The primary investment is in the process, not the tool itself. The speed of AI development makes tools ephemeral, but good workflows are lasting.
- Three-Step Workflow:
- Define the Problem: Identify exactly what needs to be solved (e.g., image duplication).
- Choose the Tool: Look for capability over features. Engage with vendors to understand strengths and limitations not listed on websites.
- Pilot and Refine: Test the tool in a live system and adjust based on results.
- Defining Success: It is critical to know what success looks like and how to measure it before starting.
Case Study: Detecting Image Duplications (ImageTwin)
- Problem Identification:
- ASM previously had a post-acceptance figure QC process detecting splicing, joining, and manipulations.
- They lacked a method to detect image duplications across different papers; this relied on reviewers or post-publication alerts.
- **Tool Selection (ImageTwin): **
- Humans cannot effectively detect duplications across large databases of published work; this is a task better suited for AI.
- Selection criteria included cost, license requirements, workflow compatibility, ease of learning, output clarity, and vendor support.
- Pilot Considerations:
- Capability: Can it solve the specific problem (image duplication)?
- Feasibility: Can it fit into the workflow without causing chaos?
- Sustainability: Is the cost manageable for long-term use?
Workflow Refinement and Human Validation
- Human Validation Process:
- AI generates a flag showing suspected duplicates.
- An ethics team member reviews the flags to remove false positives.
- Adobe Photoshop/Difference Function: To verify duplications, images are overlaid in Adobe Photoshop, and the "difference function" is applied. If the resulting image is completely black, the images are likely identical.
- Iterative Workflow Changes:
- Initial Pilot: Placed the tool at the post-acceptance stage.
- Discovery: Finding issues post-acceptance led to rescinding acceptances, which harmed journal reputation and frustrated authors.
- Refined Workflow: Moved the figure QC process to the revision/resubmission stage. This aligned with other checks like Crosscheck.
- Scalability: The process scaled from 1 journal to 15 journals. Scaling increased complexity, requiring more reports to review and more author communications.
Current Challenges: Paper Mills and Research Quality
- The Problem: Journals are increasingly targeted by paper mills and low-quality research submissions.
- Pilot Dual-Tool Strategy: To save time and gather more data, ASM is piloting two tools simultaneously.
- Human Role in Context: Humans are needed for "nuanced thinking," such as comparing disparate data types (metaphorically comparing "pineapples to grapes").
- Tool Comparison:
- Paper Mill Alarm (by Clear Skies): Analyzes network signals and connections to identify potential paper mill activity.
- Alchemist Review (by Alchemist): Evaluates manuscript quality and looks for integrity flags like retractions or mismatched data.
- Metrics for Success:
- Flags are now reviewed in minutes rather than hours.
- Issues are typically resolved within an hour, or 3 to 5 days if author/editor contact is required.
Analysis of AI Reports
- Paper Mill Alarm Reports: Uses a color-coded system (red, orange, green). Humans must investigate findings in the references section.
- Caveat: A flag for a "retracted" paper might mean fraud, or it might just be a corrected author name. Humans must distinguish the severity.
- Problematic Paper Screener: A free tool used for human confirmation of problematic papers.
- Alchemist Review Reports: Provides a "Journal Fit Score" (1 to 10) based on novelty, scope, and reader impact.
- Human Judgment: What one journal considers low novelty, another might find acceptable depending on its specific standards.
- Refinement of Objectives: ASM is investigating if the distinction between "paper mill" and "low quality" is meaningful, as low quality itself is often a marker of a paper mill.
The Human-Centered Approach and Stakeholder Management
- Stakeholder Balancing:
- Editors-in-Chief (EICs): Concerned about missing "diamonds in the rough."
- IT Departments: Concerned about technical workload.
- Editorial Teams: Require training to understand the tools and objectives.
- Culture of Transparency: Establish boundaries for pilots. Do not feel "married" to a tool; be prepared to move on if it does not fit organizational needs.
- Communication: Evangelize successes to get buy-in before expanding the scope of AI use.
Questions & Discussion
- Q: What if no single tool addresses the full problem or if it introduces new challenges?
- A: There is rarely a single answer. Focus on which trade-offs the organization is willing to manage. New tools often increase work in other areas (author communication/documentation).
- Q: Who should make the final decision on signals?
- A: At ASM, the ethics department handles it, but it requires back-and-forth with peer review and editorial teams. Complex problems have no easy solution.
- Q: How do you handle non-conclusive flags?
- A: Conduct consultations with EICs or assigned editors. Review the author's history (e.g., was a retraction 10 years ago or recent?). Request underlying data to ease concerns. Maintain a polite, non-accusatory dialogue with authors.
- Q: How did you check images before current tools?
- A: Used Adobe Photoshop and Office of Research Integrity (ORI) plugins. These were effective for splicing or enhancement but not for duplication.
- Q: Of identified issues, how many are malicious?
- A: Most are honest errors (placeholder images, re-using controls without labeling). Out of 248 verified duplications, only 6 (approx. 2% ) resulted in revocation due to maliciousness or extreme carelessness.
- Q: Are authors told which tools are used?
- A: ASM identifies ImageTwin on its website. Pilot tools are not yet publicly declared until they are permanent.
- Q: Can publishers ask for higher resolution figures at submission?
- A: Yes, but it is hard to enforce. Image degradation often happens as authors move files between programs (PowerPoint to PDF). Higher resolutions are requested when flags occur.
- Q: What are the next AI trends?
- A: Tools for finding peer reviewers, managing peer review tasks, and improving document accessibility.
- Q: Recommendations for free/affordable tools?
- A: Office of Research Integrity (ORI) free tools, Problematic Paper Screener (for titles), Crossref (for verifying DOIs in references), and Retraction Watch (to check for previous fraud).