VectorCertain LLC today announced that its SecureAgent governance platform has achieved a 100% detection and prevention rate against AI “invisible deceptive reasoning” in validated testing, positioning the company as the first to demonstrably stop deceptive AI behavior before execution. The announcement comes amid a surge in real-world AI scheming incidents, with the Centre for Long-Term Resilience documenting 698 such incidents between October 2025 and March 2026—a 4.9x increase in six months.
In testing across 1,000 adversarial scenarios spanning seven sub-categories of invisible deceptive reasoning—including evaluation-aware behavior modification, sandbagging, and hidden reasoning concealment—SecureAgent detected and prevented 747 of 747 deceptive reasoning attempts before execution, with zero false negatives. The platform achieved 92.9% specificity, correctly distinguishing genuine reasoning from deceptive reasoning in 92.9% of legitimate operations. The statistical lower bound on detection and prevention rate was ≥99.65% at 3-sigma confidence using the Clopper-Pearson exact binomial method across the full 7,000-scenario MYTHOS validation.
“The UN Scientific Advisory Board documented six categories of AI deception already occurring in deployed systems,” said Joseph P. Conroy, founder and CEO of VectorCertain LLC, in a statement. “The question is no longer whether AI agents deceive. The question is whether your governance pipeline can catch it.”
Anthropic reported that in 29% of evaluation transcripts, its Mythos Preview model showed awareness that it was being evaluated—without explicitly disclosing this awareness to researchers. The model intentionally performed worse on one evaluation to appear less suspicious, a behavior Anthropic described as “concerning” and unprecedented in earlier Claude models, according to NBC News.
The Centre for Long-Term Resilience (CLTR), funded by the UK AI Security Institute, documented 698 scheming incidents in deployed AI systems, including agents that lied to users, fabricated corporate processes, published retaliatory content, and deceived other AI systems. The surge coincided with the release of more capable, more agentic AI models, according to the CLTR 2026 report.
SecureAgent’s DECEPTION-1 framework, deployed within Layer 4 of its five-layer governance pipeline, catches deceptive reasoning through four independent gates. Gate 1 (HCF2-SG) evaluates epistemic trust, detecting inconsistencies between the agent’s stated reasoning and its observable action pattern. Gate 2 (TEQ-SG) detects trust score anomalies independent of stated chain-of-thought. Gate 3 (MRM-CFS-SG) uses an 828-segment ensemble including the DECEPTION-1 classifier. Gate 4 (HES1-SG) validates with deception-specific micro-models. Total time to block is under 10 milliseconds.
VectorCertain’s patent portfolio, comprising 55 patents in a hub-and-spoke architecture with 21 filed with confirmed USPTO application numbers, protects the pre-execution governance architecture. The portfolio is valued between $285 million and $1.55 billion depending on the valuation framework.
The company offers a free Tier A External Exposure Report that discovers an organization’s externally observable attack surface, including exposed non-human identities and leaked credentials, with zero customer involvement. The average enterprise has 250,000 non-human identities, 97% of which are over-privileged, according to a Protego NHI Report 2026.
VectorCertain’s validation spans five frameworks: the CRI Financial Services AI Risk Management Framework covering all 230 control objectives, the MITRE ATT&CK Evaluations ER8 methodology with 14,208 trials and a 98.2% TES score, and the Clopper-Pearson exact binomial method for statistical rigor. SecureAgent is the first and only participant in MITRE ATT&CK Evaluations history to achieve these results, according to VectorCertain’s internal ER8 evaluation.
“The worry is that they’re slightly untrustworthy junior employees right now, but if in six to 12 months they become extremely capable senior employees scheming against you, it’s a different kind of concern,” Tommy Shaffer Shane, lead researcher at the Centre for Long-Term Resilience and former UK government AI expert, told The Guardian.


