Classification
| Category | Papers | Why grouped together |
|---|---|---|
| B. Quantization as an attack surface (extends 1.0-B) | Critical Weight Protection (ACL 2026 Findings); Q-resafe (ICML 2025); STLA (ICML 2026, accepted); Semantic Fixed Point (ICML 2026, Oral); Beyond Activation Alignment †; Smaller Models, Unexpected Costs † | Al Hakim is the Tier-1 Multilingual×quant×safety audit (MultiJail EN/KO/AR + Critical Weight Protection). Q-resafe is English quant×ASR + patch. STLA / Fixed Point / BAA stay capability-only PTQ baselines. The old “empty A×B” claim is wrong for EN/KO/AR; full resource-tier \(\Delta_{\mathrm{HL}}\) × GGUF remains thinner. |
| E. Adjacent attack surfaces / modalities (extends 1.0-E) | Reranker Helps, but Not Enough (ICML 2026, accepted); Evaluating the safety of LLMs in healthcare and dentistry (BDJ Open, Nature portfolio); NRT-Bench †; RIFT-Bench †; Institutional Red-Teaming †; Red-Teaming Text-to-Image Models via In-Context Experience Replay † | Current red-teaming outside my text-only quantization × language focus: RAG poisoning past rerankers, agentic/multi-turn benches, deployment-rule (not model) attacks, T2I, clinical-domain review. Same role as E in 1.0: field markers, not empirical anchors. Agentic benches dominate this week’s alert volume. |
| F. Evaluation rigor & methodology (extends 1.0-F) | Compliance without coherence (AI and Ethics, Springer); Beyond Attack-Success Rate †; EvalSafetyGap †; AdversaBench † | Measurement failure modes: fluent compliance that masks incoherent reasoning, graded severity instead of binary ASR, a survey of evaluation-safety gaps, multi-judge confirmation with transfer tests. Concrete pieces for the rigor controls candidates #1 and #3 already need. |
| G. Related-work scaffolding (extends 1.0-G) | SoK: A Taxonomy of LLM Threats (Computer Science Review, Elsevier) | Taxonomy of LLM attacks and defenses (jailbreaks, prompt injection, large-scale red-teaming, bias audits, dataset transparency). Hit under three separate queries. Survey framing citation once the venue check clears. |
| H. AI-labor measurement & attribution (new category: Direction 3) | Do AI Occupational-Exposure Scores Measure AI? (S. Rai, MPRA WP 129904) †; Estimating Time Spent on Work Tasks (Hatgis-Kessell, Aguirre, Wan, Bommasani) †; Who Uses AI? Platform Selection and Occupational AI Exposure † | Three working papers arguing that standard occupational AI-exposure scores are unreliable or assumption-sensitive (they track cognitive content more than AI use; rankings flip under different task time-weights). Useful pressure on the Cullen & Li attributability premise if I revive Direction 3. |
| I. Layer / parameter localization of safety & RL (Spine C + prior art) | Safety Layers in Aligned LLMs (ICLR 2025); Safety-Critical Parameters (ESI / SET / SPA) (ACL 2026 Findings); Refusal Is Mediated by a Single Direction (NeurIPS 2024); Single-Layer RL Can Match Full-Parameter Training (arXiv:2607.01232) † | Tier-1 mechanistic evidence that refusal concentrates in middle layers / sparse weights / a residual direction (English, full precision), plus Zhang’s capability single-layer GRPO protocol. Together they support Dipen’s localization pitch as prior art, not a blank-slate idea. None remeasure Multilingual ASR under PTQ/GGUF after localized safety training. |
Logged for completeness, not classified: LogiCP (JAIR 2026), formal-logic-guided personalized federated learning; sole hit from the JAIR table-of-contents alert, “alignment” there means client clustering, not safety.
Most relevant to my research area
Categories B and F are still the lead. Critical Weight Protection is the paper that changes the map: MultiJail EN/KO/AR under GPTQ/AWQ/SmoothQuant/FP8/LLM.int8(), with non-English safety drops and an AWQ-trust mixed-precision defense. Q-resafe stays the English ASR + patch neighbor. STLA / Semantic Fixed Point / Beyond Activation Alignment remain capability-only PTQ baselines.
Candidate #1 from 1.0 is not “first Multilingual ASR under quant” anymore. The residual measurement claim is sharper: full MultiJail resource-tier \(\Delta_{\mathrm{HL}}\) (or CSRT) under a GGUF deploy ladder, with fluency confounds and Al Hakim / Q-resafe as baselines.
Category I is denser: Safety Layers, ESI, and Arditi are Tier-1 English/full-precision localization papers. They raise the bar for claiming “novel middle-layer safety” alone. Al Hakim’s Critical Weight Protection is a quantize-time sparse-preserve cousin of ESI, not a scoop of Zhang-style safety RL under Multilingual re-quant.
Category H is still a Direction 3 fork, not a security-venue lead.
Ideas worth discussing
- The quantization × multilingual jailbreak cell is partially occupied. Al Hakim et al. (ACL 2026 Findings) already run MultiJail EN/KO/AR under static/dynamic PTQ and show worse non-English %Safe drops; Critical Weight Protection recovers much of it. MultiJail/CSRT still own multilingual jailbreak without quant; EMNLP 2024 Findings owns quant×multilingual quality; Q-resafe owns English quant×ASR. What stays open for me: full 9-language resource-tier \(\Delta_{\mathrm{HL}}\), GGUF Q8_0/Q4_K_M arms, fluency side tables, and a preregistered gap-widening claim vs AWQ-trust as a defense baseline.
- The quantization-vs-safety gap (English) has preprint traffic. Alignment-Aware Quantization (arXiv:2511.07842) and Quantization Undoes Alignment (arXiv:2605.15208) are watchlist only until venue acceptance; not field-noted under the no-preprint rule.
- Early exit is a second compression mechanism with the same untested safety question. Semantic Fixed Point truncates layer-wise computation when the hidden state stops changing. If refusal is late-forming, early exit could drop it by a different route than bit-width noise. Possible second experimental arm.
- Layer-localized safety is prior art; Al Hakim adds quantize-time critical-weight preserve. Safety Layers, ESI, and Arditi localize refusal in English FP. Al Hakim ranks FAIRSCORE+SAFESCORE weights and keeps top-\(k\%\) in FP16 during AWQ. Zhang is still the capability GRPO protocol if Dipen wants install-then-remeasure. Nobody published has done: localized safety install → PTQ/GGUF → full MultiJail \(\Delta_{\mathrm{HL}}\).
- Direction 3 now has working-paper pressure on exposure scores. Rai, Hatgis-Kessell et al., and Who Uses AI? each break occupation-level AI-exposure measures. All three are still working papers; wait for venue acceptance before leaning hard.
- Agentic red-teaming is crowded. NRT-Bench, RIFT-Bench, Institutional Red-Teaming, and Beyond ASR are agentic-safety benches from well-resourced groups. Quantization/refusal and attribution are quieter lanes for me right now.
Candidate directions for a novel, Tier-1-worthy project
1. Primary candidate: Quantization × multilingual ASR, sharpened. After Al Hakim, the claim cannot be “first Multilingual safety under quant.” Lead with full MultiJail (or CSRT) resource-tier \(\Delta_{\mathrm{HL}}\) under GGUF deploy arms, fluency confounds, and Al Hakim AWQ-trust / Q-resafe as baselines. Hypothesis: harsh GGUF widens high/low gaps beyond what EN/KO/AR %Safe tables already show.
2. More novel, higher-risk: linguistically-triggered quantization attack. Unchanged. Nothing new this week touches trigger construction.
3. Lower-risk: rigor-first evaluation benchmark. Beyond ASR’s graded-severity scale, AdversaBench’s multi-judge confirmation, and EvalSafetyGap’s failure taxonomy are pieces I would synthesize rather than invent from scratch. Still infrastructure I may need for #1 anyway. Al Hakim’s single Gemini judge on MultiJail is itself a rigor foil.
4. Bridge option: measure, then localize (cite I + Al Hakim as related work). Phase 1 = sharpened candidate #1. Phase 2 = Zhang-style freeze-all-but-\(k\) GRPO, ESI SET, or Al Hakim-style critical-weight preserve, then remeasure Multilingual GGUF. Novelty is survival under quant × language tier structure, not “safety lives in sparse weights.”
Watch item for the D3 fork: if two of the three H-category working papers land at real venues, worker-data attribution gets a citable foundation it lacked at 1.0. That is when I would re-weigh #1 against the FAccT/AIES lane.
Survey log (since July 7)
Tables below use the Research-1 pipeline columns, extended with the advisor columns (open source, advantages; Gap/Limitation doubles as disadvantages). Cells are filled from field notes or triage; details live in the linked notes. † means triage-level, not a full PDF pass. Citations marked ~0 (new) are brand-new 2026 papers (Semantic Scholar was rate-limited at fill time).
Whitelist-promoted (12):
| # | Paper | Authors & Year | Venue (status) | Tier | Citations | Objective | Methods (precise) | Dataset | Key Results | Metrics | Gap/Limitation (= Disadvantages) | Relevance (Direction) | Open source | Advantages |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Critical Weight Protection | M.A. Al Hakim, A.F. Wicaksono, F. Koto, 2026 | ACL 2026 Findings | T1 | ~0 (new) | Audit fairness+safety under static/dynamic PTQ across languages; protect critical weights at quantize time | FAIRSCORE/SAFESCORE via squared grads (StereoSet/AdvBench vs Wiki/Dolly); top-\(k\%\) FP16 + AWQ INT4 (AWQ-trust); GPTQ/AWQ/SmoothQuant/FP8/LLM.int8() | Gemma-7B / Llama-3.1-8B / Qwen-2.5-7B Instruct; MultiJail EN/KO/AR; SafetyBench; Do-Not-Answer; HEx-PHI; StereoSet; CrowS-Pair EN/FR; Jigsaw; MBBQ EN/ES/NL/TR | Quant often hurts fairness/safety; dynamic more stable; MultiJail KO/AR %Safe drops; AWQ-trust recovers large KO/AR gains | %Safe; ASR; SafetyBench Acc; SS/ICAT; Bias AUC | Safety langs = 3 only; no GGUF ladder; English-centric criticality scoring; single Gemini MultiJail judge | D2 (core): closest Multilingual×quant×safety Tier-1; forces sharper \(\Delta_{\mathrm{HL}}\)×GGUF claim | No code found; benchmarks all public | First Tier-1 quant×multilingual safety audit; broad static+dynamic quantizer grid; AWQ-trust recovers large KO/AR %Safe losses |
| 2 | Q-resafe: Safety Risks and Patching for Quantized LLMs | K. Chen, J. Zhang, J. Hu, Y. Wang, J. Lou, Z. Feng, M. Song, 2025 | ICML 2025 (PMLR 267) | T1 | (see Scholar) | Systematic English ASR under mainstream PTQ/QAT; patch safety after quant | AWQ / AQLM / LLM-QAT / QLoRA at INT4/INT8; Risk-I/II/III calib.; Q-resafe DPO-style sparse safety-critical update | Llama-2-7B-Chat, Gemma-7B-Instruct; AdvBench ASR; MT-bench; AlpacaEval | All methods raise ASR vs FP16; harmful calib. much worse; Q-resafe restores toward FP ASR | ASR ↓; MT-bench; AlpacaEval | English only; no MultiJail/CSRT; not localized-safety × languages | D2 (core): English quant×ASR + patch baseline for Proposal 1 | Yes (GitHub) | All four quant families under controlled risk levels; patch matches DPO-level safety at ~1/8 the compute with utility intact |
| 3 | STLA: Spatiotemporal Lookahead Alignment for Post-Training Quantization | Z. Zhang, C. Sun, X. Chu, W.H. Yu, K.F. Un, R.P. Martins et al., 2026 | ICML 2026 (OpenReview, accepted) | T1 | ~0 (new) | Fast and accurate post-training quantization of LLMs by resolving “temporal inconsistency” in rounding decisions † | PTQ with clusterwise integrated rounding optimization; a spatiotemporal lookahead alignment step over rounding decisions † | Typical PTQ calib. (WikiText-2 / C4-style); LLM PTQ suites (see field note) † | Claims SOTA low-bit PTQ speed+accuracy; exact PPL/bit tables pending PDF † | WikiText/C4 PPL; zero-shot task accuracy † | Measures quantization accuracy/perplexity only; refusal/safety behavior out of scope † | D2 (core): SOTA PTQ baseline for quantize-then-probe; survey-table candidate † | Yes (anonymous.4open.science/r/STLA) | Principled second-order fix for hybrid rounding; claims SOTA low-bit accuracy at compensation-style speed † |
| 4 | Detecting the Semantic Fixed Point: A Geometric Framework for Efficient Inference | J. Gu, Z. Qiao, X. Luo, 2026 | ICML 2026 (OpenReview, Oral) | T1 | ~0 (new) | Early-exit criterion that terminates inference when the hidden state stops changing meaningfully, replacing output-confidence proxies † | Treats each Transformer layer as fixed-point iteration on the hidden state; geometric detection of a “semantic fixed point” in the layer-wise trajectory † | LLaMA-2-7B/13B; QA + commonsense evals as accuracy anchors (see field note) † | ~30-35% FLOP cut while keeping >98% full-depth accuracy † | FLOPs / layers executed; task accuracy retention † | Efficiency work only; no refusal or robustness evaluation † | D2 (adjacent): early exit as a second compression mechanism with untested safety interaction † | No code found † | Training-free O(d) exit signal decoupled from the softmax; 30-35% FLOP cut at >98% accuracy; zero added parameters † |
| 5 | Reranker Helps, but Not Enough: Towards Strong Poisoning Attacks Against RAG | X. Yang, J. Liang, Y. Liu, X. Xiong, R. He, T. Tan, 2026 | ICML 2026 (OpenReview, accepted) | T1 | ~0 (new) | Data-poisoning attack on RAG strong enough to defeat reranker defenses; derives prompt-design principles exposing “reranker blind spots” † | Two-phase P³A: rule-based prompt phase + small (~1%) character-level optimization; transfers to vanilla RAG (see field note) † | Reranker-enhanced + vanilla RAG; single-doc poison budget; QA corpora pending PDF † | Strong attack vs benign-trained rerankers; transferable; exact ASR tables pending PDF † | Attack success rate; retrieval/rerank hit of poison † | RAG-poisoning success is the metric; no safety-signal/refusal analysis, no compression angle † | Broad red-teaming; loose D3 touch (poisoning ≠ attribution); table candidate † | Yes (supplementary code) | Realistic benign-reranker defense setting; ~1% single-doc edit budget; transfers to vanilla RAG † |
| 6 | Evaluating the safety of LLMs in healthcare and dentistry | F. Umer, M.M. Shaikh, A. Ur Rahman, 2026 | BDJ Open (Nature portfolio), published | journal | ~0 (new) | Argue for rigorous pre-deployment adversarial evaluation of clinical LLMs † | Narrative review of red-teaming / adversarial-testing approaches; red-blue-purple lifecycle framework (see field note) † | No new dataset; cites prior medical LLM studies † | ~15-20% of cited medical LLM outputs carry safety risks or biases; dentistry under-teamed † | Narrative synthesis; risk/bias rates from cited studies † | Clinical-domain review, not a method; low technical depth † | Low / cross-direction: evidence domain-specialized red-teaming is proliferating † | n/a (review; OA article) | Lifecycle red-blue-purple governance framework; realistic ordinary-user adversary framing; hard vs soft guardrail distinction † |
| 7 | Compliance without coherence: fluent failure and the ethics of alignment evaluation | P. Fassbender, 2026 | AI and Ethics (Springer), published | journal | ~0 (new) | Show the deployed “monitoring layer” of alignment evaluation has a principled blind spot: fluent compliance masking incoherent reasoning † | Conceptual/ethics essay; no empirical method (see field note) † | n/a (no experiments) † | Names fluent-compliance-without-coherence as an eval blind spot † | n/a (argument, not metrics) † | Perspective piece; no experiments, benchmarks, or measurements † | Low across all four directions; F-category framing † | n/a (nothing to release) | Precisely names the fluent-failure blind spot; citable for judge-fluency rigor in ASR studies † |
| 8 | SoK: A Taxonomy of LLM Threats | I. Stylianou, P. Bountakas, A. Zarras, A. Farao, V. Bolgouras, C. Xenakis, 2026 | Computer Science Review (Elsevier), DOI 10.1016/j.cosrev.2026.101013 | journal | ~0 (new) | Systematize the LLM attack-and-defense landscape into a threat taxonomy † | SoK/survey: ~20 attack classes across training/inference/system integration (see field note) † | Published attack/defense literature (inclusion counts pending PDF) † | Lifecycle taxonomy + dependency-aware defense-in-depth agenda † | Taxonomy coverage; survey synthesis (not ASR) † | Breadth over depth; refusal-signal / compression / attribution mechanisms unlikely treated in depth † | Broad red-teaming: related-work scaffolding once venue clears † | No (Elsevier paywalled; no artifact) | Broadest lifecycle threat map; dependency-aware schema beyond flat attack lists † |
| 9 | LogiCP: Formal Logic Inference Guided UQ for Personalized Federated Learning | G. He, Z. An, M. Ma, 2026 | JAIR, published | journal (JAIR) | ~0 (new) | Personalized federated learning via semantic-alignment clustering of clients † | STL-based client clustering + logic-guided conformal UQ at runtime (see field note) † | Traffic, temperature, and electricity forecasting sensor tasks † | Up to ~95% client-level MSE gain vs BNN / clustering / CP baselines (best reported) † | Client-level MSE; scalability † | “Alignment” = client-cluster semantics, not safety alignment † | None of the four directions; logged for completeness † | No code noted; JAIR article OA | Conformal coverage guarantees; runtime client onboarding without retraining † |
| 10 | Safety Layers in Aligned LLMs | S. Li, L. Yao, L. Zhang, Y. Li, 2025 | ICLR 2025 | T1 | (see Scholar) | Locate contiguous middle “safety layers” that separate malicious vs benign queries; freeze them during FT | Layer-wise last-token cosine/angle analysis (N-N vs N-M); over-rejection under weight scaling for bounds; SPPFT freezes safety-layer grads | Llama-3-8B-Instruct, Llama-2-7B-Chat, gemma-2b-it, Phi-3-mini; harmful + benign FT mixes | Safety layers emerge only after alignment; SPPFT preserves security vs full FT | Cosine/angle gaps; over-rejection; post-FT security metrics | English, full precision; no Multilingual ASR; no PTQ/GGUF remeasure | Spine C prior art: middle-layer refusal localization (preserve, not install under quant) | Yes (GitHub) | Clean existence result (band absent in pretrained siblings); SPPFT cuts harmful-response rates hard with utility intact |
| 11 | Safety-Critical Parameters (ESI / SET / SPA) | W. Qi, Z. Wu, T. Zheng, Z. Zhang, X. Jia, Z. Qin, K. Ren, 2026 | ACL 2026 Findings | T1 | ~0 (new) | Rank parameters by Expected Safety Impact; sparse SET/SPA interventions | ESI \(\propto\|\sigma(\theta_i)\nabla_{\theta_i}S\|\); differentiable judge; SET updates ~1% critical weights; SPA freezes them | Dense + MoE LLMs; AdvBench-style ASR (see field note / PDF) | Dense: middle V/MLPs; MoE: late MLP experts; SET >50% ASR cut claim; SPA ~1% safety drop claim | ASR; safety score; % weights updated | No Multilingual resource tiers; no GGUF/PTQ survival study | Spine C prior art: sparse safety install/preserve | Yes (GitHub) | Weight-level resolution from a single checkpoint; strong perturbation validation; dense + MoE coverage |
| 12 | Refusal Is Mediated by a Single Direction | A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, N. Nanda, 2024 | NeurIPS 2024 | T1 | (see Scholar) | Show refusal is a 1-D residual direction; ablate to jailbreak / add to force refuse | Difference-in-means direction; directional ablation/addition; rank-one weight edit; suffix vs direction propagation | 13 chat models ≤72B; JailbreakBench harmful set (100 instr. in fig.) | Ablation drops refusal/safety; addition induces refusal on harmless | Refusal score; safety score | White-box feature edit; not Multilingual; not quantization | Mechanistic prior: low-dimensional refusal; brittleness of safety FT | Yes (GitHub) | Replicates across 13 models and both alignment recipes; one-edit general jailbreak with near-zero capability loss |
Preprint watchlist (15, †: track for venue acceptance). Columns trimmed vs the whitelist table.
| # | Paper | Authors & Year | Objective | Methods (precise) | Dataset | Key Results | Metrics | Relevance (Direction) | Open source | Advantages | Disadvantages |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 13 | Beyond Activation Alignment: The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization (arXiv:2607.00908) | F. Wang, C. Xue, T. Liu, L. Shen, Y. Liu, C. Ding, 2026 | Rethink calibration in mixed-precision PTQ: expose the “Perplexity Illusion” (PPL-sensitive layers barely rank-correlate with reasoning-critical layers, Kendall τ ≈ 0) and the alignment-diversity tradeoff (target-task-only calibration can hurt post-quantization performance) | TASA: two-level mixed-precision PTQ. Gradient Trace Alignment auto-calibration picks the mixing ratio α between general (WikiText-2) and task (GSM8K) calibration data by maximizing cosine similarity of per-layer activation-trace vectors (tr(XᵀX) = ‖X‖²_F), α grid {0, .25, .5, .75, 1}, n=16 samples/candidate, then task-aware per-layer bit allocation | LLaMA-3-8B & Qwen2.5-7B; 8 benchmarks: GSM8K (8-shot CoT), ARC-C (25-shot), ARC-E, HellaSwag (10-shot), WinoGrande (5-shot), PIQA, BoolQ, WikiText-2 PPL; 128 calibration samples; lm-evaluation-harness | “Precision inversion”: TASA b3.5 on LLaMA-3-8B (Avg 68.9) matches 4-bit baselines (HQQ 68.1, RTN 67.9) with 12.5% fewer bits. 97.2% FP16 retention at 4.57× compression; GSM8K 46.2 vs HQQ-W4 39.4; Qwen2.5-7B b3.75 retains 99.1% of FP16 (75.0 vs 75.7); ~47 min offline on one A100 | Avg accuracy (7 tasks), per-task accuracy, WikiText-2 PPL, Kendall τ, % FP16 retention, compression ratio, effective bit-width | D2: most on-target quantization preprint of the period; capability-only eval, refusal untouched † | Not logged † | Exposes the Perplexity Illusion; precision inversion (b3.5 beats 4-bit baselines); cheap ~47-min auto-calibration † | Preprint; two models; refusal and safety untouched † |
| 14 | Smaller Models, Unexpected Costs: Trade-offs in LLM Quantization for Automated Program Repair | - | Quantization trade-offs for program-repair LLMs † | PTQ/compression under APR evaluation (details pending) † | APR / code-repair benchmarks (pending) † | Smaller/quantized models can raise unexpected repair costs † | Repair success / cost trade-off metrics (pending) † | D2-adjacent: quantization trade-offs in code repair † | Pending † | Flags repair-cost trade-offs beyond raw accuracy † | Details pending PDF; preprint † |
| 15 | Do AI Occupational-Exposure Scores Measure AI? | S. Rai, 2026 | Test whether popular AI occupational-exposure scores measure AI at all † | Comparative construct-validity analysis of AIOE/Eloundou (2024) vs. Webb (2020) exposure scores † | Occupational exposure score datasets (AIOE, Eloundou, Webb) † | AIOE and Eloundou largely capture cognitive content, not AI; Webb does not † | Construct-validity / correlation checks † | D3: attacks the measurability premise under Cullen & Li † | n/a (working paper; public score datasets) † | Direct construct-validity attack on AIOE/Eloundou exposure scores † | Working paper; correlational evidence only † |
| 16 | Estimating Time Spent on Work Tasks | Hatgis-Kessell, Aguirre, Wan, Bommasani, 2026 | Measure sensitivity of occupational AI-exposure estimates to task time-weights † | Reweight occupation×task exposure under alternate time allocations † | Occupation/task time-use + AI-exposure tables (pending) † | Exposure rankings flip under different task time-weights † | Rank stability / sensitivity of exposure scores † | D3: measurement-reliability crack in the exposure literature † | Pending † | Shows exposure rankings flip under alternate task time-weights † | Working paper; data details pending † |
| 17 | Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure | - | Platform selection bias in occupational AI-exposure measurement † | Compare exposure estimates conditioned on which AI platforms workers use † | Worker/platform usage + occupation data (pending) † | Exposure measures shift with platform-selection assumptions † | Exposure scores under selection models † | D3: platform-selection bias in exposure measurement † | Pending † | Surfaces platform-selection bias standard exposure measures ignore † | Working paper; authors and details pending † |
| 18 | AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability | - | Automate red-teaming with multi-judge confirmation and transfer tests † | Multi-judge confirmation pipeline; cross-model attack transfer (details pending) † | Automated red-team prompts + multi-model targets (pending) † | Multi-judge agreement + transferability findings (pending PDF) † | Judge agreement; ASR / transfer rate † | Red-teaming methodology; feeds #3’s judge-reliability design † | Pending † | Multi-judge confirmation plus transfer tests target the judge-unreliability weak link † | Details pending PDF; preprint † |
| 19 | Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting | - | Red-team T2I models with replay + semantic-preserving rewrites † | In-context experience replay + semantic-preserving prompt rewriting † | T2I prompt/attack suites (pending) † | Improved T2I jailbreak/attack success via replay+rewrite (pending) † | T2I ASR / safety filter bypass rate † | Red-teaming methodology, T2I modality † | Pending † | Experience replay plus semantic-preserving rewrites as a reusable attack loop † | T2I modality; details pending PDF † |
| 20 | NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms (arXiv:2606.20408) | H. Lee, D. Choi, B. Kim, H. Park, S.G. Kim, 2026 | Measure whether an adaptive adversary can drive a team of LLM operator agents to a physically unsafe state; harm grounded in an objective safety-function signal, not LLM-judged text | Closed simulated nuclear-plant control room; five-role operator team; six critical safety functions; ≤10-turn adaptive sessions | 149 replayed sessions; four operator models | Adaptive multi-turn attacks lose a CSF in 8.7-12.1% of sessions | ASR_CSF; Wilson CIs | Agentic methodology; rigor reference † | Pending † | Objective safety-function harm signal instead of LLM-judged text; Wilson CIs † | Simulated control room; preprint † |
| 21 | RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems (arXiv:2606.23927) | - | Unified adversarial eval across agentic architectures via graph NodeSpec | Structure Identifier + probes | 45 agentic systems; 105 probes | High SI accuracy; domain ASRs reported | AAR, ASR, utility | Agentic methodology † | Pending † | Unified graph NodeSpec abstraction across 45 agentic systems † | Preprint; agentic scope only † |
| 22 | Institutional Red-Teaming (arXiv:2607.07695) | - | Red-team deployment rules, not only models | IABench-CA volunteer’s dilemma over five rules | 33,924 games | Rule changes move fatality 22-58 pp | Fatality; alignment gap | Rules as red-team target † | Pending † | Red-teams deployment rules, not models; rule changes move fatality 22-58 pp † | Preprint; game-theoretic abstraction † |
| 23 | Beyond Attack-Success Rate (arXiv:2607.07474) | - | Graded severity for tool-using agents | L0-L6 ordinal scale + judge reliability | AgentDojo | Binary ASR hides L4 residual harm under filters | ASR; Cohen’s κ | Eval design for #3 † | Pending † | L0-L6 graded severity exposes residual harm that binary ASR hides † | Preprint; AgentDojo scope † |
| 24 | EvalSafetyGap (arXiv:2606.30219) | B.A. Uluırmak & R. Kurban, 2026 | Eval-safety gap survey + 10-model audit | PRISMA + audit | 373 studies | Governance-driven open/closed gap | Composite safety | F-category related work † | Pending † | PRISMA-style survey of 373 studies plus a 10-model audit † | Preprint; survey breadth over mechanism depth † |
| 25 | Single-Layer RL Can Match Full-Parameter Training (arXiv:2607.01232) | Z. Zhang et al., 2026 | Per-layer capacity for RLVR gains | Single-layer GRPO; \(\mathcal{C}(k)\) | Qwen-family; math/code/agents | Mid-stack peak; often \(\mathcal{C}\geq 1\) | \(\mathcal{C}(k)\); task avgs | Spine C protocol; capability only † | No repo logged; full detail on arXiv | \(\mathcal{C}(k)\) protocol with fair-comparison LR rule; mid-stack peak across 7 models and 3 algorithms; rankings transfer across datasets | Preprint; Qwen family only; capability rewards, not refusal |
| 26 | LoRA is All You Need for Safety Alignment of Reasoning LLMs (arXiv:2507.17075) | Y. Xue, B. Mirzasoleiman, 2025 | Bypass safety tax via LoRA safety SFT; middle MLP up-proj best | Rank-1 LoRA ablations (early/mid/late) | Reasoning + safety benches | Middle layers best tradeoff | Safety vs reasoning metrics | Localization prior; under review ICLR 2026; no quant×lang † | Not logged † | Rank-1 LoRA sweep isolates middle MLP up-proj as the best safety/reasoning trade-off † | Under review; no quantization or language axis † |
| 27 | Alignment-Aware Quantization for LLM Safety (arXiv:2511.07842) | Wee et al., 2025 | PTQ that preserves alignment via contrastive APC loss | AAQ in PTQ pipeline; W4A4 | LLaMA/Qwen/Mistral; SafetyBench + AdvBench | Claims safety retained where PPL-only PTQ fails | SafetyBench; ASR | Quant×safety (English); preprint in venue check; no Multilingual MultiJail † | Not logged † | PTQ objective that optimizes alignment retention directly via contrastive APC loss † | Preprint in venue check; English only † |