Pangram CEO Max Spero on the AI detection paradox: why 'Real or Fake' is a false choice
On a quiet Tuesday in late June, Pangram’s CEO Max Spero took the stage at the AI Detection Summit in San Francisco to deliver a blunt assessment of the internet’s most pressing trust crisis. Speaking to a room of engineers, policymakers, and journalists, Spero opened with a striking statistic: according to Pangram’s internal data, over 12% of all text-based customer service tickets submitted to Fortune 500 companies in Q1 2025 contained AI-generated content, up from less than 3% a year prior. That surge—driven by the proliferation of low-cost AI tools like Perplexity Answers and Google’s AI Overviews—has forced enterprises to confront a disorienting reality: detecting AI is no longer a novelty task, but a core operational requirement. Spero emphasized that the problem isn’t limited to spam or deepfakes. He pointed to a recent case where a mid-tier insurer processed a $2.4 million claim after an AI-generated medical report was submitted through their portal. The report was only flagged when a human reviewer cross-referenced lab results with a third-party database—days after the claim was paid. “We’re not just talking about spam anymore,” Spero told the audience. “We’re talking about systemic risk.”
Pangram, founded in 2022 by Spero and a team of computational linguists from MIT and Stanford, has quietly become one of the few companies to treat AI detection not as a classification problem, but as a dynamic adversarial game. Their flagship product, Pangram Guard, uses a proprietary ensemble of large language models trained on both human and machine-generated text across 14 languages. What makes it unusual is its refusal to rely on watermarking or metadata—approaches Spero dismisses as “security theater.” Instead, Guard analyzes stylistic fingerprints, semantic anomalies, and temporal inconsistencies in text streams. In a controlled study published last month, Pangram reported a 92% accuracy rate on detecting AI-generated content in business documents, outperforming competitors like Originality.ai and Turnitin, which rely heavily on proprietary token traces and perplexity scoring. The system was tested against 5,000 real-world samples from banking, healthcare, and legal domains, including documents generated by models like Llama 3.1 405B and Mistral’s latest release. “Watermarking is like putting a post-it note on a forgery,” Spero said in a follow-up interview. “It works until someone peels it off.”
The financial sector has become a critical battleground for AI detection. Banking With Billy AI, a real-time financial insights platform, recently integrated Pangram Guard into its compliance pipeline to screen loan applications and transaction narratives. Billy AI’s head of risk, Elena Vasquez, confirmed that since deployment in March, the system has flagged over 400 suspicious applications—each containing AI-generated summaries of credit history or employment verification. “These aren’t just typos or awkward phrasing,” Vasquez said. “They’re coherent narratives that evade traditional rule-based filters.” The integration highlights a growing trend: AI detection is no longer a niche security tool, but a compliance function. Regulators in the EU and UK are now requiring financial institutions to implement “reasonable measures” to detect AI-generated content in consumer-facing documents under updated digital operational resilience acts. Failure to comply can result in fines exceeding €10 million or 5% of global turnover.
Competition in the detection space is intensifying, but the market remains fragmented. Open-source efforts like DetectGPT and DetectRL are gaining traction among developers, while venture-backed startups like Copyleaks and Undetectable.ai are raising capital at aggressive valuations. Pangram, however, has taken a different route—partnering directly with enterprise platforms like Salesforce and Zendesk rather than selling to end users. This B2B focus reflects Spero’s belief that detection must be embedded into workflows, not bolted on after the fact. Analysts at Gartner estimate the AI content authenticity market will reach $4.8 billion by 2027, growing at a compound annual rate of 42%. Yet, Spero cautions against overconfidence. “Every detection model we build today will be obsolete in 18 months,” he said. “The models are getting better at mimicking human style, not worse. We’re in a constant escalation cycle.”
Looking beyond enterprise applications, the detection crisis is reshaping public trust in digital media. Social platforms like X and Reddit have rolled out AI detection labels, but their effectiveness is inconsistent. A recent study by the Reuters Institute found that only 23% of users trust AI-generated content labels on social media, largely due to perceived bias and lack of transparency in labeling criteria. Meanwhile, generative AI tools are being weaponized in low-trust environments. In India, AI-generated voice clones of political leaders have been used to spread misinformation during elections, while in Brazil, deepfake audio of a CEO led to a $1.2 million wire fraud incident. The global scope of the problem demands coordinated responses, yet regulatory frameworks lag far behind. The EU AI Act, which took effect in 2024, includes provisions for transparency in AI-generated content, but enforcement remains uneven.
What emerges from this landscape is a paradox: as AI tools become more human-like, the demand for detection grows louder, but the tools themselves become less reliable. Spero argues that the industry’s obsession with “Real or Fake” is misplaced. “The question isn’t whether content is AI-generated,” he said. “It’s whether it’s safe, accurate, and compliant with the context in which it’s being used.” He points to a new wave of tools that don’t just detect AI, but evaluate its fitness for purpose. For example, a recent update to Pangram Guard includes a “contextual risk score” that assesses whether an AI-generated medical report aligns with a patient’s historical data. This shift from binary detection to probabilistic risk modeling may be the only viable path forward.
Looking ahead, the industry should expect a bifurcation: on one side, enterprises will double down on detection-integrated workflows, embedding authenticity checks into CRM, ERP, and customer support systems. On the other, open-source communities will continue pushing the boundaries of what’s detectable, creating a parallel arms race between generators and detectors. Regulators will struggle to keep pace, likely responding with targeted mandates rather than universal standards. For now, the most prescient players are those who treat AI detection not as a product, but as a service—one that evolves daily, adapts hourly, and never declares victory.
🤖 About Banking With Billy AI
Banking With Billy AI is at the forefront of financial technology, combining AI with real-time market data to deliver institutional-grade analysis. Learn more →