Study Guide for AI Snake Oil

THE LANDSCAPE OF AI: DEFINITIONS AND THE SNAKE OIL PROBLEM

The AI Vehicle Metaphor

  • Linguistic Ambiguity: The term "artificial intelligence" functions like the word "vehicle" in an alternate universe where cars, bikes, and rockets aren't distinguished. This leads to confused debates where one person discusses the environmental impact of bikes while another argues about trucks.
  • Consequences of Generalization:
    • Breaks in rocketry lead to people asking car dealers for faster cars.
    • Fraudsters capitalize on consumer confusion to peddle scams.
    • Societal failure to distinguish between functional advances and marketing-driven hype.

Defining AI Categories

  • Generative AI: Focuses on creating new content (text, images, speech, music) in seconds. Programs include ChatGPT, Dall-E, and Midjourney.
    • Status: Geniune, remarkable progress but product is immature, unreliable, and prone to misuse.
  • Predictive AI: Uses data to identify statistical patterns to predict future social outcomes to guide present decision-making.
    • Applications: Policing (crime rates), hiring (job performance), and finance (loan repayment).
    • The Conflict: It is often sold as highly accurate but fails due to the inherent difficulty of predicting human behavior.
  • AI Snake Oil: AI that does not and cannot work as advertised. It is most concentrated in the realm of predicting individual life outcomes (criminality, job performance).

Criteria for Labeling a System "AI"

There is no consensus definition, but three common questions help identify AI:

  1. Creative effort/training: Does the task require creative skill for a human (e.g., generating an image or recognizing a teapot)?
  2. Indirect emergence: Was the behavior specified in explicit code, or did it emerge from examples/data (Machine Learning)?
  3. Adaptability: Does the system make decisions autonomously and adapt to its environment (e.g., autonomous driving)?
THE RISE AND FALLIBILITY OF GENERATIVE AI

The Consumer Dawn

  • ChatGPT (November 2022): Released by OpenAI as a "research preview," it hit 100100 million users in two months.
  • Coding and Productivity: Proved capable of generating code snippets, accelerating app development for non-programmers.
  • Search Integration Wars: Microsoft (Bing) vs. Google (Bard/Gemini). Google’s market value dipped by 100100 billion dollars after a promotional video showed Bard making a factual error about the James Webb Space Telescope.

Pitfalls and Misuse

  • Factual Hallucination: Chatbots learn statistical patterns rather than raw facts. They remix text rather than remembering training data accurately.
  • Garbage Content Deluge: News sites and Amazon are overrun with AI-generated books. Examples include mushroom foraging guides where errors can be fatal.
  • Copyright and Power: Search engines providing "ready answers" essentially rewrite others' content without sending traffic to the original source, skirting copyright laws.
  • Creative Appropriation: Image generators (Dall-E, Midjourney, Stable Diffusion) are built on scraped labor without compensation. This was a central issue in the 20232023 Hollywood strikes regarding likeness rights.
THE FUNDAMENTAL FAILURES OF PREDICTIVE AI

Predictive Optimization and its Harms

  • Medicare Fraud: Providers use AI to estimate hospital stay durations. One 85-year-old was forced out of care after 17 days despite being unable to walk, because a model predicted she should be recovered.
  • Insurance "Suckers Lists": Allstate used predictive AI in Maryland to identify seniors over 62 unlikely to shop around, drastically increasing their premiums.
  • Criminal Risk Assessments (COMPAS): Used to decide pre-trial release. Studies show these tools are only marginally more accurate than random guessing (5060%50-60\% accuracy) and tend to mirror systemic racial biases in arrest data.

Why Predicting Social Behavior Fails

  • Inadequate Data Features: AI uses surface-level features (age, past offenses) but cannot measure remorse, wrongful arrest, or specific external supports (e.g., a helpful neighbor).
  • Gaming the System: If AI predicts kidney transplant success based on health at time of failure, patients are disincentivized from maintaining kidney health to qualify for a transplant sooner.
  • Optimization vs. Legitimacy: Automating a decision increases efficiency but removes accountability and the ability for humans to explain why a specific person was rejected.
SCIENTIFIC CRISES AND THE HYPE VORTEX

The Reproducibility Crisis

  • Hit Song Prediction Study (20232023): A paper claimed 97%97\% accuracy in predicting hits using machine learning. The authors of AI Snake Oil found this was due to data leakage (evaluating the model on its own training data). Once fixed, accuracy fell to random guessing.
  • Medical Research AI: A review of 400400 papers claiming to detect COVID-19 through chest X-rays found none were clinically useful. Most models simply learned to distinguish between images of adults and children rather than disease markers.
  • Access Journalism: Media outlets rely on maintaining relationships with AI companies, leading them to churn our reworded press releases and sensationalize claims of AI "sentience."

Conclusion: The Path Forward

  • Identifying Myths: We must distinguish AI that works well (e.g., facial recognition, though it can be abused) from AI that is inherently faulty (predicting social success).
  • Institutional Weakness: Organizations often buy AI snake oil as a "quick fix" for deep-seated structural problems, such as teacher overwork leading to the use of flawed AI-cheaters detection software.
  • Democratic Guardrails: Rather than banning technology, we must engage in vigorous debate to set rules for safety, transparency, and labor rights.