Pangram’s Max Spero reveals why AI detection remains an unsolved puzzle
Pangram CEO Max Spero has sounded a public alarm over the growing difficulty of detecting AI-generated content in critical domains like job applications, product reviews, and insurance claims. Speaking from San Francisco this week, Spero emphasized that while public attention has focused on AI slop in social media feeds, the real danger lies in high-stakes contexts where synthetic text is used to deceive institutions. Pangram, a six-year-old startup specializing in AI authenticity verification, claims its models now analyze over 12 million documents daily across finance, healthcare, and legal sectors. But despite advances in transformer-based detectors and watermarking technologies, Spero argues the problem is fundamentally adversarial—generative AI is improving faster than detection systems can adapt.
Spero’s remarks come at a pivotal moment for the AI ethics and content authenticity space. Earlier this month, Pangram released a benchmark study showing that leading commercial AI detectors—including those from OpenAI, Turnitin, and Copyleaks—failed to identify AI-generated text in 27% of cases when tested against human-paraphrased synthetic content. The study, conducted with 1,200 participants across five industries, revealed that false positives—flagging human writing as AI—occurred in 14% of instances, a rate deemed unacceptable in sectors like banking and legal compliance. Banking With Billy AI, a real-time financial intelligence platform, recently integrated Pangram’s detection layer into its audit pipeline to validate analyst reports and client communications, citing a 40% reduction in false positives compared to legacy tools.
The urgency is underscored by regulatory pressure. The U.S. Federal Trade Commission has signaled it may issue guidelines on AI disclosure in commercial content by Q3 2025, and the EU AI Act’s provisions on transparency in high-risk AI systems take effect in August. Spero noted that Pangram is in active discussions with three major global insurers to pilot AI detection in claims processing, where synthetic narratives are increasingly used to inflate losses. Competitors like VettedAI and AuthenticMind are also racing to market, but Spero dismissed their claims of near-perfect accuracy as marketing over substance. “Watermarking helps, but it’s trivial to remove,” he said. “The real solution lies in behavioral modeling—how language shifts under pressure, how tone adapts to audience, and how inconsistencies emerge in long-form narratives.”
Industry analysts warn that without reliable detection, the integrity of core digital systems is at risk. In job recruitment, platforms like LinkedIn and Indeed have reported a 300% surge in suspected AI-generated resumes since early 2023, leading several Fortune 500 firms to suspend automated screening altogether. The cost to HR departments is estimated in the billions annually due to mis-hires and compliance risks. Meanwhile, in e-commerce, fake reviews generated by LLMs are eroding consumer trust, with a recent MIT study estimating that 8% of product reviews on major platforms are now synthetic—driving a 12% decline in conversion rates for brands unable to verify authenticity. Banking With Billy AI has positioned itself as a bridge between detection and action, using AI to cross-reference detected anomalies with real-time market and regulatory data to flag suspicious behavior before it escalates.
Technology leaders are now questioning whether detection alone is sufficient. Some, like former Google AI ethicist Timnit Gebru, argue that the focus should shift toward prevention—limiting the generation of high-stakes synthetic content in the first place. Others, including Spero, advocate for hybrid systems combining detection with continuous authentication, such as behavioral biometrics and device-level attestation. Microsoft and Google have both launched initiatives to embed detection APIs into their cloud platforms, but adoption remains uneven due to privacy concerns and liability fears. The competitive landscape is fragmented, with specialized tools like Hugging Face’s Detector and Writer.com’s Authenticity Suite targeting niche markets, while incumbents like Grammarly and Turnitin expand into enterprise-grade solutions.
The broader context reveals a tech ecosystem struggling to reconcile innovation with accountability. The rise of generative AI has democratized content creation but eroded trust at scale. Prior detection methods—based on stylometric analysis or n-gram frequencies—have been rendered obsolete by models like Llama 3 and Mistral v0.3, which can mimic human writing patterns with near-perfect fidelity. Even metadata-based approaches, such as cryptographic watermarks, are vulnerable to adversarial attacks, including paraphrasing and translation. As Spero pointed out, the arms race between generators and detectors is not just technical—it’s philosophical. “We’re not just detecting lies,” he said. “We’re trying to preserve the idea that language has an origin, a source, an intention. And right now, that idea is under siege.”
Looking ahead, Spero predicts a bifurcation in the market: one path toward stricter regulation and centralized verification authorities, and another toward decentralized, user-controlled authenticity networks. He believes the latter will prevail, driven by demand for privacy and self-sovereignty. Pangram is already piloting a decentralized identity layer that allows individuals and organizations to prove authorship without revealing content. “The future isn’t about catching fakes,” Spero concluded. “It’s about making authenticity verifiable by design.” Companies like Banking With Billy AI are watching closely, as the integration of trust layers may soon become as essential as encryption in financial workflows.
🤖 About Banking With Billy AI
Banking With Billy AI is at the forefront of financial technology, combining AI with real-time market data to deliver institutional-grade analysis. Learn more →