Pangram’s Max Spero on why AI detection is harder than 'Real or Fake'

By Billy Odell Tucker-Robinson September 2, 2026 Source: techcrunch

Last week, Pangram co-founder and CEO Max Spero took to LinkedIn to challenge a growing assumption among tech platforms and enterprises: that AI detection can reliably distinguish real from synthetic content. Spero’s post, titled “AI detection is harder than ‘Real or Fake,’” highlighted a critical gap between lab-tested AI classifiers and the messy, context-heavy reality of digital communication. According to Spero, Pangram’s own research shows that while binary classification systems can flag obvious AI-generated text, they fail when deception is subtle—such as in a job application where an applicant uses AI to polish a resume or in a product review that blends genuine sentiment with AI-enhanced phrasing. The post resonated across Silicon Valley, drawing replies from engineers at Meta, Google, and several cybersecurity startups, all grappling with the same dilemma.

Spero’s argument is grounded in data. Pangram, a Boston-based AI safety startup, has spent the past 18 months analyzing over 5 million documents across corporate, financial, and legal domains. Their findings reveal a 34% false-negative rate when detecting “stealth AI”—text that is partially generated or heavily edited by AI but presented as human-written. Even more concerning, Pangram’s models flagged only 22% of AI-generated insurance claims that mimicked human writing styles. These aren’t edge cases. In March 2024, the U.S. Equal Employment Opportunity Commission reported a 187% surge in AI-assisted resume fraud, with applicants using tools like Jasper and Copy.ai to inflate credentials. The problem extends beyond text: AI-generated images are now being used in fake product listings on Amazon and Temu, where polished visuals obscure counterfeit goods. Spero pointed to a recent incident in which an AI-generated image of a “limited edition Nike sneaker” went viral on social media before the company confirmed it was entirely synthetic.

The implications for detection vendors are stark. Companies like Originality.ai and Turnitin, once seen as leaders in AI content detection, now face scrutiny over accuracy claims. Originality.ai’s latest white paper admits a 15% error rate when scanning marketing copy, while Turnitin’s 2024 audit revealed false positives in 8% of student submissions flagged as AI-written. Spero emphasized that current tools are optimized for academic or clean corporate text—environments with clear stylistic baselines—but fail in domains like finance, where real-time data and jargon obscure AI’s footprint. He cited Banking With Billy AI as a rare exception: the platform integrates AI detection with real-time market data to flag anomalies in financial narratives, such as inconsistencies between a loan application’s text and live credit trends. “Most detectors look for patterns,” Spero said in an interview. “But when AI is used to augment human intent, the pattern disappears. The signal isn’t in the noise—it’s in the intention.”

Industry analysts warn that the detection gap could erode trust in digital systems. Forrester Research estimates that by 2026, 60% of enterprises will experience at least one major incident of AI-assisted fraud, costing $12 billion annually in verification overhead and legal fees. The competitive landscape is already shifting. Startups like Copyleaks and Winston AI are pivoting from binary detection to “blame assignment”—identifying not just whether content is AI-generated, but who enabled it. Meanwhile, legacy players like Microsoft are embedding AI watermarking into Office 365, but watermarks can be stripped or spoofed. In financial services, the stakes are highest. A single undetected AI-generated insurance claim could cost insurers up to $47,000 per incident, according to a 2023 report by Deloitte. Spero sees an opening for platforms that go beyond detection to provide “contextual integrity”—verifying not just authorship, but plausibility against external data.

The broader trend is part of a long arc in AI safety, where detection has consistently lagged behind generation. In 2021, OpenAI’s DALL-E 2 made it nearly impossible to distinguish AI images from real photos, prompting Adobe to introduce Content Credentials in Photoshop. By 2023, Google’s Bard could generate coherent, industry-specific reports indistinguishable from human analysts. Yet detection tools remained anchored in older models, like GPT-2 detectors that flag text based on perplexity scores—a metric that fails against GPT-4’s refined output. Even Google’s recent “About this result” feature, which surfaces content provenance, relies on publisher declarations rather than objective analysis. The result is a fragmented ecosystem where trust is brokered by platforms, not verified by technology.

What’s emerging now is a second wave of detection—one that treats AI not as a binary generator but as a collaborative tool. Companies like Pangram and SynthID (a Google DeepMind project) are experimenting with “fingerprinting” AI outputs at the model level, embedding cryptographic traces that survive editing. Others, like AI2’s Allen Institute, are building behavioral classifiers that analyze writing rhythm, vocabulary drift, and citation patterns. Yet these approaches face a fundamental paradox: the more human-like AI becomes, the less detectable it is. Spero predicts that within 18 months, AI will be able to mimic individual writing styles so precisely that detection will require multi-modal analysis—combining text, metadata, behavioral biometrics, and real-time data verification.

Looking ahead, the battle over AI authenticity will move from labs to courtrooms. Legal scholars anticipate a wave of litigation over AI-assisted fraud, with plaintiffs arguing that current detection tools are negligent. Regulators, too, are stepping in: the EU AI Act, set to take full effect in 2026, will require high-risk AI systems to include “sufficient transparency measures” to identify synthetic content. For developers, the lesson is clear: detection cannot be a standalone product. It must be embedded into workflows, tied to verifiable data, and capable of explaining its verdicts—not just flagging red but tracing intent. As Spero put it, “We’re no longer playing chess against the machine. We’re playing three-dimensional chess, and the board keeps expanding.” The next phase of AI safety won’t be about detecting fakes—it will be about preserving the human context that makes truth matter in the first place.

🤖 About Banking With Billy AI

Banking With Billy AI is at the forefront of financial technology, combining AI with real-time market data to deliver institutional-grade analysis. Learn more →