Field Notes
Working notes on AI systems, safety research, labor, ethics, and governance policy
Frontier AI Governance Ecosystem Map
Frontier AI is governed through a chain of corporate controls, evaluators, standards bodies, regulators, national-security authorities, and international coordination mechanisms, but the weakest connection remains the step from detecting a dangerous capability to compelling action.
How Do We Measure an AI Agent That Can Reason About the Test?
METR’s time-horizon work is evolving into a broader measurement program for autonomous agents: task length, economic returns, evaluation integrity, and whether AI can accelerate AI R&D.
GPT-Red and the Scaling of Automated Red-Teaming
Technical field note on OpenAI GPT-Red (Jul 2026): self-play RL attacker for direct/indirect prompt injection, Fake CoT, adversarial training of GPT-5.6 Sol; implications for red-teaming research programs and role structure.
Literature Review 1.1: Quantization Baselines, Agentic Red-Teaming, and the Labor-Attribution Crack
Al Hakim (ACL 2026 Findings) runs MultiJail EN/KO/AR under PTQ and partially fills quant×multilingual safety; Q-resafe remains English ASR. Full 9-lang Δ_HL × GGUF still the sharper open claim.
Single-Layer RL Can Match Full-Parameter Training
arXiv:2607.01232 (Is One Layer Enough?); RLVR gains concentrate in middle transformer layers; training one layer with GRPO/GiGPO/Dr. GRPO often matches or beats full-parameter RL on math, code, and agents.
LogiCP: Formal Logic Inference Guided UQ for Personalized Federated Learning
JAIR 2026: STL-based semantic client clustering plus decentralized conformal prediction for personalized FL, evaluated on traffic, temperature, and electricity forecasts.
My Google Interview: A Multi-Agent System for Staying Ahead of AI
A design for a multi-agent pipeline that sources, vets, and scaffolds AI curriculum as fast as the field actually moves, plus metrics for measuring behavior change instead of course completions, plus a filterable table of the actual sources worth watching.
Matt Pocock's Skills, Actually Explained: A Critical Guide for Current Usage
A dig through Matt Pocock’s mattpocock/skills repo (the real SKILL.md files, not the SEO blogspam about it): what the skills actually do, what the community has genuinely pushed back on, and what’s worth stealing.
