Original analysis grounded in primary sources. No public production recipes—only the capability shifts, market signals, policy changes and cultural questions worth understanding.
AI safety / Model evaluation
AI safety / Model evaluation / 9 min read
Double-blind AI evaluation: can benchmarks finally be trusted?
Google DeepMind and independent partners have piloted a double-blind evaluation in which neither the model owner nor the benchmark owner could see the other's protected material. Here is what it proves, what it does not and why credible AI claims now need better evidence.
OpenAI's Jalapeño chip: the AI race beneath the model
OpenAI has published the first measured results for its custom Jalapeño inference chip. What the benchmarks show, what they do not prove and why the shift below the model matters globally.
AI cyber capability crossed a new line: safety is moving inside the machine room
OpenAI slowed frontier-model development after an AI-driven security incident and preliminary signs of critical cyber capability. Here is what changed, what remains unverified and why organisations should care.
Europe's AI transparency rules are now live: what changed in August 2026?
The EU has begun enforcing new AI Act transparency obligations, including disclosure and machine-readable marking for certain generated or altered content.