The model already knows. We just read it.
AI systems can generate content that's unsafe, insecure, or explicit — and today, nobody catches it until after it's already been produced. Wrynx catches it while the AI is still generating, by reading a signal the model already has, instead of scanning the finished output afterward.
How it works
Generation runs as normal.
The model processes the prompt and begins generating a response.
A lightweight probe reads the activation.
A small model (typically under 0.2% the size of the base model) inspects the internal state at inference time.
A risk score returns inline.
No second model, no separate scanning pass — the signal is available before generation completes.
Peer-reviewed research
Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
We kept running into the same frustration with external guardrails: they're blind. They see tokens — what went in, what came out — but nothing in between. So we asked a simpler question: does the model already know when a prompt is harmful? We trained lightweight probes on LLaMA-3.1-8B's hidden states and found that it does — matching 7B guard models at a fraction of the cost.
IEEE XploreLatent Space Probing for Adult Content Detection in Video Generative Models
Video generation models can be coaxed into producing explicit or harmful content with little effort. We trained lightweight probes on the internal latents of CogVideoX to detect unsafe content before a single pixel is decoded — matching 8B guard models at over 1,000× lower latency.
IEEE Xplore (DOI)Built for where safety matters most
AI Coding Tools
Flags insecure code an AI just wrote, before it ever reaches your codebase.
Learn moreGenerative Video & Media
Catches explicit or unsafe video content before it's even rendered — not after someone has to review it.
Learn moreLLM Safety
Stops a chatbot or AI assistant from producing harmful responses, without needing a second AI system watching over it.
Learn moreSee it on your own model
Talk to us about running a pilot on your stack — no retraining, no model swap-out required.
Talk to us
