As “AI-washing” spreads across the compliance sector, Theta Lake is making the case that the difference between genuine AI-native platforms and superficial marketing claims will define which RegTech vendors survive the next wave of workplace technology governance.
Theta Lake recently discussed the evolution of machine learning at the company and building a multimodal foundation.
The communications compliance firm argues its credentials are structural rather than cosmetic. Its first corporate hire was a chief data scientist, its compliance classifiers have used AI since inception, and its architecture rests on proprietary patents dating back to 2018.
The company, named a Visionary in the Gartner Magic Quadrant and holding independently verified ISO 42001 and CSA Star for AI Level 2 certifications, has now lifted the lid on how its machine learning architecture evolved, with distinguished engineer Rohit Jain detailing the journey.
From the outset, Theta Lake was designed for a multimodal world. The firm recognised that legacy lexicon-based systems built for email could not cope with the nuance of modern channels such as video and chat.
Lexicons cannot generalise beyond syntax into semantics and suffer from high false positive rates because they lack context awareness, a problem that has intensified as new communication modalities reshape language itself.
The company’s first hurdle was moving beyond rigid keyword detection. Rather than forcing customers to abandon existing lexicons, it engineered proprietary IP (US Patent No. 12,045,561) to handle lookalikes and soundalikes, keeping the system robust against misspellings, OCR errors and semantic variations.
Further patented work extended fuzzy matching across larger sections of text, with corresponding error models for OCR and transcription.
Data quality proved equally critical. Theta Lake rigorously tested multiple third-party OCR and transcription engines across varied audio and video conditions, speakers and accents, producing a bespoke transcription test suite it runs periodically to benchmark and select the highest-performing vendors.
On the modelling side, the firm progressed from word embeddings to sentence embeddings, selecting those with the optimal balance of sensitivity and specificity on its datasets.
These embeddings feed ensembles of traditional machine learning methods, spanning Naive Bayes, neural networks, tree-based methods, boosting, GBMs and KNeighbors, in a process that has moved from manual to almost completely automated. Where accuracy gains justify the performance overhead, fine-tuned discriminative and generative language models are deployed selectively.
Every classifier the company builds today is a custom-tuned ensemble, and the same data-driven philosophy governs its visual pipeline, which has expanded from object detection to image classification, including patented methods for detecting applications shared on screen.
Read the full Theta Lake post here.
Copyright © 2026 FinTech Global









