Field Notes
Working notes on AI systems, safety research, labor, ethics, and governance policy
A Longitudinal Database of Agentic-AI Security Incidents
54 publicly disclosed incidents from Feb 2023 to Sep 2026, coded on one schema: the count of demonstrations grows steadily, but what changes is that 21 of them were seen in the wild, none before 2025.
Frontier AI Governance Ecosystem Map
Frontier AI is governed through a chain of corporate controls, evaluators, standards bodies, regulators, national-security authorities, and international coordination mechanisms, but the weakest connection remains the step from detecting a dangerous capability to compelling action.
How Do We Measure an AI Agent That Can Reason About the Test?
METR’s time-horizon work is evolving into a broader measurement program for autonomous agents: task length, economic returns, evaluation integrity, and whether AI can accelerate AI R&D.
GPT-Red and the Scaling of Automated Red-Teaming
Technical field note on OpenAI GPT-Red (Jul 2026): self-play RL attacker for direct/indirect prompt injection, Fake CoT, adversarial training of GPT-5.6 Sol; implications for red-teaming research programs and role structure.
Literature Review 1.1: Quantization Baselines, Agentic Red-Teaming, and the Labor-Attribution Crack
Al Hakim (ACL 2026 Findings) runs MultiJail EN/KO/AR under PTQ and partially fills quant×multilingual safety; Q-resafe remains English ASR. Full 9-lang Δ_HL × GGUF still the sharper open claim.
Single-Layer RL Can Match Full-Parameter Training
arXiv:2607.01232 (Is One Layer Enough?); RLVR gains concentrate in middle transformer layers; training one layer with GRPO/GiGPO/Dr. GRPO often matches or beats full-parameter RL on math, code, and agents.
LogiCP: Formal Logic Inference Guided UQ for Personalized Federated Learning
JAIR 2026: STL-based semantic client clustering plus decentralized conformal prediction for personalized FL, evaluated on traffic, temperature, and electricity forecasts.
My Google Interview: A Multi-Agent System for Staying Ahead of AI
A design for a multi-agent pipeline that sources, vets, and scaffolds AI curriculum as fast as the field actually moves, plus metrics for measuring behavior change instead of course completions, plus a filterable table of the actual sources worth watching.
Matt Pocock's Skills, Actually Explained: A Critical Guide for Current Usage
A dig through Matt Pocock’s mattpocock/skills repo (the real SKILL.md files, not the SEO blogspam about it): what the skills actually do, what the community has genuinely pushed back on, and what’s worth stealing.
AI Capabilities Are Accelerating Faster Than Our Ability to Measure Their Risks
The International AI Safety Report 2026 describes a widening gap between rapidly improving AI capabilities and our ability to evaluate, govern, and build resilience around them. The central problem may increasingly be measurement itself.
