Writing
Practical thinking on shipping AI that's production-ready and provable: what breaks, how to catch it, and how to prove it to your team, your customers, and your auditors. No hype. No vendor comparisons.
AI Security Reviews Are Blocking Your Deals. Here's the Evidence Buyers Actually Want
Your AI feature works. Then an enterprise prospect sends a vendor security questionnaire with a new AI section, and the deal stalls for six weeks while you assemble evidence by hand. Here's what reviewers are actually asking for, and how to have it ready.
How to Monitor LLM Calls in Production: A Complete Setup Guide
Standard infrastructure monitoring tells you the service is up. It doesn't tell you whether the model is producing correct outputs, whether latency is acceptable at p95, or whether costs are tracking. Here's the complete setup: what to instrument, what metrics to track, and which tools to use.
SOC 2 for AI Workflows: What an Auditor Actually Checks
SOC 2 wasn't written with LLMs in mind, but auditors are now applying its Trust Services Criteria to AI systems anyway. Here's how the criteria map to the reality of an AI workflow, and the control gaps that surface most often.
The Real Cost of Running Unmonitored AI in Production
The team ships the AI feature. It works in staging. Production looks clean. But the outputs have been wrong at 12% since launch. Costs are running 3x the estimate. Nobody knows yet. This is the unmonitored AI problem, and this post quantifies what it actually costs.
The OWASP LLM Top 10, Translated for People Who Have to Sign Off on It
The OWASP Top 10 for LLM Applications is the closest thing to a shared security standard for AI. But it's written for engineers. Here's what each risk means for the person who has to attest that the AI is safe, and the one question to ask about each.
Six Gaps Even a Good Eval Suite Usually Has
Most eval suites have happy-path bias, don't block deployment, lack regression testing, and go stale. An eval suite with these gaps can report 94% accuracy while missing a 15% failure rate on real production inputs. Here are the six gaps we see most often, and what to add.
Where Personal Data Actually Leaks in LLM Apps (GDPR & HIPAA)
Most teams building AI features have a privacy policy and encrypted databases. Then an LLM call quietly ships a customer's personal data to a third-party provider that may retain it. Here's where personal data actually leaks, and what GDPR and HIPAA require you to do about it.
Production RAG: What Nobody Tells You After 6 Months
The RAG tutorials get you to a demo in an afternoon. They don't cover what happens six months into production: index staleness, retrieval quality decay, RAG-specific hallucination modes, cost at scale, and the chunking strategy that made sense at launch but doesn't fit real usage. Here's what we've learned.
The EU AI Act Is Here. A Practical Readiness Checklist for Engineering Teams
The EU AI Act is the first broad AI regulation with real teeth, and its obligations are phasing in. You don't need to become a lawyer, but you do need to know your risk tier and have the technical evidence ready. Here's a practical checklist for engineering teams.
See how your AI workflows actually score.
Production and compliance readiness, from one scan — 115+ production controls across 9 dimensions, plus a compliance assessment against the EU AI Act and ISO/IEC 42001. Free for your first workflow. No credit card required.
Scan Your Repo, Free →