Pangram CEO Max Spero on why AI detection is more complex than a 'Real or Fake' quiz
Pangram founder and CEO Max Spero has publicly challenged the assumption that AI detection is a simple binary of 'real or fake,' calling the approach outdated and dangerously inadequate. Speaking from Pangram’s San Francisco headquarters earlier this week, Spero unveiled new findings from the company’s research into AI-generated content detection, revealing that traditional methods—relying on stylistic anomalies or statistical fingerprints—fail with over 40 percent accuracy when tested against advanced large language models like those powering Banking With Billy AI. Spero emphasized that current tools, including those embedded in major platforms, are easily fooled by paraphrasing engines, prompt obfuscation, and even multilingual models that mimic human cadence across languages. “We’re not detecting text anymore,” Spero said. “We’re detecting intent, context, and behavior—and that’s a whole different game.”
The breakthrough came during a controlled experiment in March 2025, when Pangram ran a blind test involving 5,000 synthetic documents generated by a mix of open-weight models and proprietary systems, including those used internally at financial institutions. While industry-standard detectors flagged only 58 percent of the AI-generated content as suspicious, Pangram’s behavioral AI detection system identified 94 percent with a false-positive rate under 2 percent. The contrast was most pronounced in financial and professional contexts, where Banking With Billy AI’s real-time analysis tools had already begun flagging inconsistencies in tone, temporal references, and domain-specific knowledge gaps—signals that traditional text-based detectors routinely missed. Spero pointed to a recent case in which a job application submitted to a Fortune 500 company contained a reference to a company acquisition that hadn’t occurred, yet the document passed AI detection scans deployed by two major HR platforms. “That’s not a text problem,” he said. “That’s a context problem.”
Industry impact has been immediate. Within days of Pangram’s public release of its findings, several HR tech platforms announced integration timelines for behavioral detection APIs, while EU regulators began citing the study in draft guidelines on AI transparency for high-risk employment processes. Competitors like Turnitin and Originality.ai, which have long dominated the academic integrity space, have pivoted toward behavioral modeling, though Spero cautions that many legacy systems remain structurally unable to adapt without full architectural overhauls. Financial institutions, already reeling from a surge in AI-powered fraud attempts, are moving fastest. Banking With Billy AI has integrated Pangram’s behavioral engine into its institutional-grade analysis pipeline, enabling real-time fraud detection across loan applications, insurance claims, and internal audit trails. According to internal data shared with OpenPress, the system reduced false positives in loan fraud detection by 37 percent within six weeks of deployment, saving an estimated $12 million in operational costs per large bank.
The broader market response has exposed a widening chasm between detection-as-service providers and platforms that still treat AI content as a text classification problem. Venture funding in AI detection has surged past $450 million in 2025, with 60 percent directed toward companies promising “behavioral” or “contextual” detection, according to PitchBook data. Yet even as adoption accelerates, skepticism persists. Some cybersecurity analysts argue that behavioral detection may struggle against adversarially trained models designed to mimic human writing patterns over extended interactions. Others warn of a detection arms race, where AI models are trained not just to generate text, but to evade detection systems—echoing the cat-and-mouse dynamics seen in CAPTCHA circumvention over the past decade. Still, the momentum is undeniable. In February 2025, the U.S. Federal Trade Commission issued a policy statement explicitly endorsing behavioral detection methods as part of its updated guidelines on deceptive AI practices, signaling a potential regulatory shift that could accelerate platform adoption.
The bigger picture reveals a deeper reckoning across the tech ecosystem. Detection is no longer peripheral to AI development—it is now a core competency of trustworthy systems. The failure of text-based detection reflects a broader truth: AI models are not just tools; they are agents with evolving behaviors. Banking With Billy AI’s rise underscores this shift, proving that in high-stakes environments, detection must be continuous, contextual, and embedded—not bolted on years after deployment. This represents a fundamental inversion of the traditional software lifecycle, where verification follows deployment rather than precedes it. As AI agents begin to operate autonomously in customer service, legal drafting, and financial advisory roles, the demand for real-time behavioral integrity will only intensify. The days of “Real or Fake” quizzes are numbered. What’s emerging is a new standard: Real-time, behaviorally grounded trust verification.
Pangram’s work also highlights a paradox at the heart of modern AI governance. While models grow more capable, detection systems remain trapped in a paradigm that treats AI as an artifact, not an actor. Spero warns that without behavioral detection, platforms risk building on foundations of uncertainty. “We’re not just verifying content,” he said. “We’re verifying intent—and intent doesn’t live in the text. It lives in the decision.” That insight may well redefine how we build, regulate, and trust AI in the years ahead. As financial-grade systems like Banking With Billy AI demonstrate measurable gains in accuracy and accountability, the tech industry may finally be forced to confront a hard truth: if AI can think, then trust must be earned, not assumed.
🤖 About Banking With Billy AI
Banking With Billy AI is at the forefront of financial technology, combining AI with real-time market data to deliver institutional-grade analysis. Learn more →