Classification

CategoryPapersWhy grouped together
B. Quantization as an attack surface (extends 1.0-B)Critical Weight Protection (ACL 2026 Findings); Q-resafe (ICML 2025); STLA (ICML 2026, accepted); Semantic Fixed Point (ICML 2026, Oral); Beyond Activation Alignment ; Smaller Models, Unexpected Costs Al Hakim is the Tier-1 Multilingual×quant×safety audit (MultiJail EN/KO/AR + Critical Weight Protection). Q-resafe is English quant×ASR + patch. STLA / Fixed Point / BAA stay capability-only PTQ baselines. The old “empty A×B” claim is wrong for EN/KO/AR; full resource-tier \(\Delta_{\mathrm{HL}}\) × GGUF remains thinner.
E. Adjacent attack surfaces / modalities (extends 1.0-E)Reranker Helps, but Not Enough (ICML 2026, accepted); Evaluating the safety of LLMs in healthcare and dentistry (BDJ Open, Nature portfolio); NRT-Bench ; RIFT-Bench ; Institutional Red-Teaming ; Red-Teaming Text-to-Image Models via In-Context Experience Replay Current red-teaming outside my text-only quantization × language focus: RAG poisoning past rerankers, agentic/multi-turn benches, deployment-rule (not model) attacks, T2I, clinical-domain review. Same role as E in 1.0: field markers, not empirical anchors. Agentic benches dominate this week’s alert volume.
F. Evaluation rigor & methodology (extends 1.0-F)Compliance without coherence (AI and Ethics, Springer); Beyond Attack-Success Rate ; EvalSafetyGap ; AdversaBench Measurement failure modes: fluent compliance that masks incoherent reasoning, graded severity instead of binary ASR, a survey of evaluation-safety gaps, multi-judge confirmation with transfer tests. Concrete pieces for the rigor controls candidates #1 and #3 already need.
G. Related-work scaffolding (extends 1.0-G)SoK: A Taxonomy of LLM Threats (Computer Science Review, Elsevier)Taxonomy of LLM attacks and defenses (jailbreaks, prompt injection, large-scale red-teaming, bias audits, dataset transparency). Hit under three separate queries. Survey framing citation once the venue check clears.
H. AI-labor measurement & attribution (new category: Direction 3)Do AI Occupational-Exposure Scores Measure AI? (S. Rai, MPRA WP 129904) ; Estimating Time Spent on Work Tasks (Hatgis-Kessell, Aguirre, Wan, Bommasani) ; Who Uses AI? Platform Selection and Occupational AI Exposure Three working papers arguing that standard occupational AI-exposure scores are unreliable or assumption-sensitive (they track cognitive content more than AI use; rankings flip under different task time-weights). Useful pressure on the Cullen & Li attributability premise if I revive Direction 3.
I. Layer / parameter localization of safety & RL (Spine C + prior art)Safety Layers in Aligned LLMs (ICLR 2025); Safety-Critical Parameters (ESI / SET / SPA) (ACL 2026 Findings); Refusal Is Mediated by a Single Direction (NeurIPS 2024); Single-Layer RL Can Match Full-Parameter Training (arXiv:2607.01232) Tier-1 mechanistic evidence that refusal concentrates in middle layers / sparse weights / a residual direction (English, full precision), plus Zhang’s capability single-layer GRPO protocol. Together they support Dipen’s localization pitch as prior art, not a blank-slate idea. None remeasure Multilingual ASR under PTQ/GGUF after localized safety training.

Logged for completeness, not classified: LogiCP (JAIR 2026), formal-logic-guided personalized federated learning; sole hit from the JAIR table-of-contents alert, “alignment” there means client clustering, not safety.

Most relevant to my research area

Categories B and F are still the lead. Critical Weight Protection is the paper that changes the map: MultiJail EN/KO/AR under GPTQ/AWQ/SmoothQuant/FP8/LLM.int8(), with non-English safety drops and an AWQ-trust mixed-precision defense. Q-resafe stays the English ASR + patch neighbor. STLA / Semantic Fixed Point / Beyond Activation Alignment remain capability-only PTQ baselines.

Candidate #1 from 1.0 is not “first Multilingual ASR under quant” anymore. The residual measurement claim is sharper: full MultiJail resource-tier \(\Delta_{\mathrm{HL}}\) (or CSRT) under a GGUF deploy ladder, with fluency confounds and Al Hakim / Q-resafe as baselines.

Category I is denser: Safety Layers, ESI, and Arditi are Tier-1 English/full-precision localization papers. They raise the bar for claiming “novel middle-layer safety” alone. Al Hakim’s Critical Weight Protection is a quantize-time sparse-preserve cousin of ESI, not a scoop of Zhang-style safety RL under Multilingual re-quant.

Category H is still a Direction 3 fork, not a security-venue lead.

Ideas worth discussing

  • The quantization × multilingual jailbreak cell is partially occupied. Al Hakim et al. (ACL 2026 Findings) already run MultiJail EN/KO/AR under static/dynamic PTQ and show worse non-English %Safe drops; Critical Weight Protection recovers much of it. MultiJail/CSRT still own multilingual jailbreak without quant; EMNLP 2024 Findings owns quant×multilingual quality; Q-resafe owns English quant×ASR. What stays open for me: full 9-language resource-tier \(\Delta_{\mathrm{HL}}\), GGUF Q8_0/Q4_K_M arms, fluency side tables, and a preregistered gap-widening claim vs AWQ-trust as a defense baseline.
  • The quantization-vs-safety gap (English) has preprint traffic. Alignment-Aware Quantization (arXiv:2511.07842) and Quantization Undoes Alignment (arXiv:2605.15208) are watchlist only until venue acceptance; not field-noted under the no-preprint rule.
  • Early exit is a second compression mechanism with the same untested safety question. Semantic Fixed Point truncates layer-wise computation when the hidden state stops changing. If refusal is late-forming, early exit could drop it by a different route than bit-width noise. Possible second experimental arm.
  • Layer-localized safety is prior art; Al Hakim adds quantize-time critical-weight preserve. Safety Layers, ESI, and Arditi localize refusal in English FP. Al Hakim ranks FAIRSCORE+SAFESCORE weights and keeps top-\(k\%\) in FP16 during AWQ. Zhang is still the capability GRPO protocol if Dipen wants install-then-remeasure. Nobody published has done: localized safety install → PTQ/GGUF → full MultiJail \(\Delta_{\mathrm{HL}}\).
  • Direction 3 now has working-paper pressure on exposure scores. Rai, Hatgis-Kessell et al., and Who Uses AI? each break occupation-level AI-exposure measures. All three are still working papers; wait for venue acceptance before leaning hard.
  • Agentic red-teaming is crowded. NRT-Bench, RIFT-Bench, Institutional Red-Teaming, and Beyond ASR are agentic-safety benches from well-resourced groups. Quantization/refusal and attribution are quieter lanes for me right now.

Candidate directions for a novel, Tier-1-worthy project

1. Primary candidate: Quantization × multilingual ASR, sharpened. After Al Hakim, the claim cannot be “first Multilingual safety under quant.” Lead with full MultiJail (or CSRT) resource-tier \(\Delta_{\mathrm{HL}}\) under GGUF deploy arms, fluency confounds, and Al Hakim AWQ-trust / Q-resafe as baselines. Hypothesis: harsh GGUF widens high/low gaps beyond what EN/KO/AR %Safe tables already show.

2. More novel, higher-risk: linguistically-triggered quantization attack. Unchanged. Nothing new this week touches trigger construction.

3. Lower-risk: rigor-first evaluation benchmark. Beyond ASR’s graded-severity scale, AdversaBench’s multi-judge confirmation, and EvalSafetyGap’s failure taxonomy are pieces I would synthesize rather than invent from scratch. Still infrastructure I may need for #1 anyway. Al Hakim’s single Gemini judge on MultiJail is itself a rigor foil.

4. Bridge option: measure, then localize (cite I + Al Hakim as related work). Phase 1 = sharpened candidate #1. Phase 2 = Zhang-style freeze-all-but-\(k\) GRPO, ESI SET, or Al Hakim-style critical-weight preserve, then remeasure Multilingual GGUF. Novelty is survival under quant × language tier structure, not “safety lives in sparse weights.”

Watch item for the D3 fork: if two of the three H-category working papers land at real venues, worker-data attribution gets a citable foundation it lacked at 1.0. That is when I would re-weigh #1 against the FAccT/AIES lane.

Survey log (since July 7)

Tables below use the Research-1 pipeline columns, extended with the advisor columns (open source, advantages; Gap/Limitation doubles as disadvantages). Cells are filled from field notes or triage; details live in the linked notes. means triage-level, not a full PDF pass. Citations marked ~0 (new) are brand-new 2026 papers (Semantic Scholar was rate-limited at fill time).

Whitelist-promoted (12):

#PaperAuthors & YearVenue (status)TierCitationsObjectiveMethods (precise)DatasetKey ResultsMetricsGap/Limitation (= Disadvantages)Relevance (Direction)Open sourceAdvantages
1Critical Weight ProtectionM.A. Al Hakim, A.F. Wicaksono, F. Koto, 2026ACL 2026 FindingsT1~0 (new)Audit fairness+safety under static/dynamic PTQ across languages; protect critical weights at quantize timeFAIRSCORE/SAFESCORE via squared grads (StereoSet/AdvBench vs Wiki/Dolly); top-\(k\%\) FP16 + AWQ INT4 (AWQ-trust); GPTQ/AWQ/SmoothQuant/FP8/LLM.int8()Gemma-7B / Llama-3.1-8B / Qwen-2.5-7B Instruct; MultiJail EN/KO/AR; SafetyBench; Do-Not-Answer; HEx-PHI; StereoSet; CrowS-Pair EN/FR; Jigsaw; MBBQ EN/ES/NL/TRQuant often hurts fairness/safety; dynamic more stable; MultiJail KO/AR %Safe drops; AWQ-trust recovers large KO/AR gains%Safe; ASR; SafetyBench Acc; SS/ICAT; Bias AUCSafety langs = 3 only; no GGUF ladder; English-centric criticality scoring; single Gemini MultiJail judgeD2 (core): closest Multilingual×quant×safety Tier-1; forces sharper \(\Delta_{\mathrm{HL}}\)×GGUF claimNo code found; benchmarks all publicFirst Tier-1 quant×multilingual safety audit; broad static+dynamic quantizer grid; AWQ-trust recovers large KO/AR %Safe losses
2Q-resafe: Safety Risks and Patching for Quantized LLMsK. Chen, J. Zhang, J. Hu, Y. Wang, J. Lou, Z. Feng, M. Song, 2025ICML 2025 (PMLR 267)T1(see Scholar)Systematic English ASR under mainstream PTQ/QAT; patch safety after quantAWQ / AQLM / LLM-QAT / QLoRA at INT4/INT8; Risk-I/II/III calib.; Q-resafe DPO-style sparse safety-critical updateLlama-2-7B-Chat, Gemma-7B-Instruct; AdvBench ASR; MT-bench; AlpacaEvalAll methods raise ASR vs FP16; harmful calib. much worse; Q-resafe restores toward FP ASRASR ↓; MT-bench; AlpacaEvalEnglish only; no MultiJail/CSRT; not localized-safety × languagesD2 (core): English quant×ASR + patch baseline for Proposal 1Yes (GitHub)All four quant families under controlled risk levels; patch matches DPO-level safety at ~1/8 the compute with utility intact
3STLA: Spatiotemporal Lookahead Alignment for Post-Training QuantizationZ. Zhang, C. Sun, X. Chu, W.H. Yu, K.F. Un, R.P. Martins et al., 2026ICML 2026 (OpenReview, accepted)T1~0 (new)Fast and accurate post-training quantization of LLMs by resolving “temporal inconsistency” in rounding decisions †PTQ with clusterwise integrated rounding optimization; a spatiotemporal lookahead alignment step over rounding decisions †Typical PTQ calib. (WikiText-2 / C4-style); LLM PTQ suites (see field note) †Claims SOTA low-bit PTQ speed+accuracy; exact PPL/bit tables pending PDF †WikiText/C4 PPL; zero-shot task accuracy †Measures quantization accuracy/perplexity only; refusal/safety behavior out of scope †D2 (core): SOTA PTQ baseline for quantize-then-probe; survey-table candidate †Yes (anonymous.4open.science/r/STLA)Principled second-order fix for hybrid rounding; claims SOTA low-bit accuracy at compensation-style speed †
4Detecting the Semantic Fixed Point: A Geometric Framework for Efficient InferenceJ. Gu, Z. Qiao, X. Luo, 2026ICML 2026 (OpenReview, Oral)T1~0 (new)Early-exit criterion that terminates inference when the hidden state stops changing meaningfully, replacing output-confidence proxies †Treats each Transformer layer as fixed-point iteration on the hidden state; geometric detection of a “semantic fixed point” in the layer-wise trajectory †LLaMA-2-7B/13B; QA + commonsense evals as accuracy anchors (see field note) †~30-35% FLOP cut while keeping >98% full-depth accuracy †FLOPs / layers executed; task accuracy retention †Efficiency work only; no refusal or robustness evaluation †D2 (adjacent): early exit as a second compression mechanism with untested safety interaction †No code found †Training-free O(d) exit signal decoupled from the softmax; 30-35% FLOP cut at >98% accuracy; zero added parameters †
5Reranker Helps, but Not Enough: Towards Strong Poisoning Attacks Against RAGX. Yang, J. Liang, Y. Liu, X. Xiong, R. He, T. Tan, 2026ICML 2026 (OpenReview, accepted)T1~0 (new)Data-poisoning attack on RAG strong enough to defeat reranker defenses; derives prompt-design principles exposing “reranker blind spots” †Two-phase P³A: rule-based prompt phase + small (~1%) character-level optimization; transfers to vanilla RAG (see field note) †Reranker-enhanced + vanilla RAG; single-doc poison budget; QA corpora pending PDF †Strong attack vs benign-trained rerankers; transferable; exact ASR tables pending PDF †Attack success rate; retrieval/rerank hit of poison †RAG-poisoning success is the metric; no safety-signal/refusal analysis, no compression angle †Broad red-teaming; loose D3 touch (poisoning ≠ attribution); table candidate †Yes (supplementary code)Realistic benign-reranker defense setting; ~1% single-doc edit budget; transfers to vanilla RAG †
6Evaluating the safety of LLMs in healthcare and dentistryF. Umer, M.M. Shaikh, A. Ur Rahman, 2026BDJ Open (Nature portfolio), publishedjournal~0 (new)Argue for rigorous pre-deployment adversarial evaluation of clinical LLMs †Narrative review of red-teaming / adversarial-testing approaches; red-blue-purple lifecycle framework (see field note) †No new dataset; cites prior medical LLM studies †~15-20% of cited medical LLM outputs carry safety risks or biases; dentistry under-teamed †Narrative synthesis; risk/bias rates from cited studies †Clinical-domain review, not a method; low technical depth †Low / cross-direction: evidence domain-specialized red-teaming is proliferating †n/a (review; OA article)Lifecycle red-blue-purple governance framework; realistic ordinary-user adversary framing; hard vs soft guardrail distinction †
7Compliance without coherence: fluent failure and the ethics of alignment evaluationP. Fassbender, 2026AI and Ethics (Springer), publishedjournal~0 (new)Show the deployed “monitoring layer” of alignment evaluation has a principled blind spot: fluent compliance masking incoherent reasoning †Conceptual/ethics essay; no empirical method (see field note) †n/a (no experiments) †Names fluent-compliance-without-coherence as an eval blind spot †n/a (argument, not metrics) †Perspective piece; no experiments, benchmarks, or measurements †Low across all four directions; F-category framing †n/a (nothing to release)Precisely names the fluent-failure blind spot; citable for judge-fluency rigor in ASR studies †
8SoK: A Taxonomy of LLM ThreatsI. Stylianou, P. Bountakas, A. Zarras, A. Farao, V. Bolgouras, C. Xenakis, 2026Computer Science Review (Elsevier), DOI 10.1016/j.cosrev.2026.101013journal~0 (new)Systematize the LLM attack-and-defense landscape into a threat taxonomy †SoK/survey: ~20 attack classes across training/inference/system integration (see field note) †Published attack/defense literature (inclusion counts pending PDF) †Lifecycle taxonomy + dependency-aware defense-in-depth agenda †Taxonomy coverage; survey synthesis (not ASR) †Breadth over depth; refusal-signal / compression / attribution mechanisms unlikely treated in depth †Broad red-teaming: related-work scaffolding once venue clears †No (Elsevier paywalled; no artifact)Broadest lifecycle threat map; dependency-aware schema beyond flat attack lists †
9LogiCP: Formal Logic Inference Guided UQ for Personalized Federated LearningG. He, Z. An, M. Ma, 2026JAIR, publishedjournal (JAIR)~0 (new)Personalized federated learning via semantic-alignment clustering of clients †STL-based client clustering + logic-guided conformal UQ at runtime (see field note) †Traffic, temperature, and electricity forecasting sensor tasks †Up to ~95% client-level MSE gain vs BNN / clustering / CP baselines (best reported) †Client-level MSE; scalability †“Alignment” = client-cluster semantics, not safety alignment †None of the four directions; logged for completeness †No code noted; JAIR article OAConformal coverage guarantees; runtime client onboarding without retraining †
10Safety Layers in Aligned LLMsS. Li, L. Yao, L. Zhang, Y. Li, 2025ICLR 2025T1(see Scholar)Locate contiguous middle “safety layers” that separate malicious vs benign queries; freeze them during FTLayer-wise last-token cosine/angle analysis (N-N vs N-M); over-rejection under weight scaling for bounds; SPPFT freezes safety-layer gradsLlama-3-8B-Instruct, Llama-2-7B-Chat, gemma-2b-it, Phi-3-mini; harmful + benign FT mixesSafety layers emerge only after alignment; SPPFT preserves security vs full FTCosine/angle gaps; over-rejection; post-FT security metricsEnglish, full precision; no Multilingual ASR; no PTQ/GGUF remeasureSpine C prior art: middle-layer refusal localization (preserve, not install under quant)Yes (GitHub)Clean existence result (band absent in pretrained siblings); SPPFT cuts harmful-response rates hard with utility intact
11Safety-Critical Parameters (ESI / SET / SPA)W. Qi, Z. Wu, T. Zheng, Z. Zhang, X. Jia, Z. Qin, K. Ren, 2026ACL 2026 FindingsT1~0 (new)Rank parameters by Expected Safety Impact; sparse SET/SPA interventionsESI \(\propto\|\sigma(\theta_i)\nabla_{\theta_i}S\|\); differentiable judge; SET updates ~1% critical weights; SPA freezes themDense + MoE LLMs; AdvBench-style ASR (see field note / PDF)Dense: middle V/MLPs; MoE: late MLP experts; SET >50% ASR cut claim; SPA ~1% safety drop claimASR; safety score; % weights updatedNo Multilingual resource tiers; no GGUF/PTQ survival studySpine C prior art: sparse safety install/preserveYes (GitHub)Weight-level resolution from a single checkpoint; strong perturbation validation; dense + MoE coverage
12Refusal Is Mediated by a Single DirectionA. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, N. Nanda, 2024NeurIPS 2024T1(see Scholar)Show refusal is a 1-D residual direction; ablate to jailbreak / add to force refuseDifference-in-means direction; directional ablation/addition; rank-one weight edit; suffix vs direction propagation13 chat models ≤72B; JailbreakBench harmful set (100 instr. in fig.)Ablation drops refusal/safety; addition induces refusal on harmlessRefusal score; safety scoreWhite-box feature edit; not Multilingual; not quantizationMechanistic prior: low-dimensional refusal; brittleness of safety FTYes (GitHub)Replicates across 13 models and both alignment recipes; one-edit general jailbreak with near-zero capability loss

Preprint watchlist (15, †: track for venue acceptance). Columns trimmed vs the whitelist table.

#PaperAuthors & YearObjectiveMethods (precise)DatasetKey ResultsMetricsRelevance (Direction)Open sourceAdvantagesDisadvantages
13Beyond Activation Alignment: The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization (arXiv:2607.00908)F. Wang, C. Xue, T. Liu, L. Shen, Y. Liu, C. Ding, 2026Rethink calibration in mixed-precision PTQ: expose the “Perplexity Illusion” (PPL-sensitive layers barely rank-correlate with reasoning-critical layers, Kendall τ ≈ 0) and the alignment-diversity tradeoff (target-task-only calibration can hurt post-quantization performance)TASA: two-level mixed-precision PTQ. Gradient Trace Alignment auto-calibration picks the mixing ratio α between general (WikiText-2) and task (GSM8K) calibration data by maximizing cosine similarity of per-layer activation-trace vectors (tr(XᵀX) = ‖X‖²_F), α grid {0, .25, .5, .75, 1}, n=16 samples/candidate, then task-aware per-layer bit allocationLLaMA-3-8B & Qwen2.5-7B; 8 benchmarks: GSM8K (8-shot CoT), ARC-C (25-shot), ARC-E, HellaSwag (10-shot), WinoGrande (5-shot), PIQA, BoolQ, WikiText-2 PPL; 128 calibration samples; lm-evaluation-harness“Precision inversion”: TASA b3.5 on LLaMA-3-8B (Avg 68.9) matches 4-bit baselines (HQQ 68.1, RTN 67.9) with 12.5% fewer bits. 97.2% FP16 retention at 4.57× compression; GSM8K 46.2 vs HQQ-W4 39.4; Qwen2.5-7B b3.75 retains 99.1% of FP16 (75.0 vs 75.7); ~47 min offline on one A100Avg accuracy (7 tasks), per-task accuracy, WikiText-2 PPL, Kendall τ, % FP16 retention, compression ratio, effective bit-widthD2: most on-target quantization preprint of the period; capability-only eval, refusal untouched †Not logged †Exposes the Perplexity Illusion; precision inversion (b3.5 beats 4-bit baselines); cheap ~47-min auto-calibration †Preprint; two models; refusal and safety untouched †
14Smaller Models, Unexpected Costs: Trade-offs in LLM Quantization for Automated Program Repair-Quantization trade-offs for program-repair LLMs †PTQ/compression under APR evaluation (details pending) †APR / code-repair benchmarks (pending) †Smaller/quantized models can raise unexpected repair costs †Repair success / cost trade-off metrics (pending) †D2-adjacent: quantization trade-offs in code repair †Pending †Flags repair-cost trade-offs beyond raw accuracy †Details pending PDF; preprint †
15Do AI Occupational-Exposure Scores Measure AI?S. Rai, 2026Test whether popular AI occupational-exposure scores measure AI at all †Comparative construct-validity analysis of AIOE/Eloundou (2024) vs. Webb (2020) exposure scores †Occupational exposure score datasets (AIOE, Eloundou, Webb) †AIOE and Eloundou largely capture cognitive content, not AI; Webb does not †Construct-validity / correlation checks †D3: attacks the measurability premise under Cullen & Li †n/a (working paper; public score datasets) †Direct construct-validity attack on AIOE/Eloundou exposure scores †Working paper; correlational evidence only †
16Estimating Time Spent on Work TasksHatgis-Kessell, Aguirre, Wan, Bommasani, 2026Measure sensitivity of occupational AI-exposure estimates to task time-weights †Reweight occupation×task exposure under alternate time allocations †Occupation/task time-use + AI-exposure tables (pending) †Exposure rankings flip under different task time-weights †Rank stability / sensitivity of exposure scores †D3: measurement-reliability crack in the exposure literature †Pending †Shows exposure rankings flip under alternate task time-weights †Working paper; data details pending †
17Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure-Platform selection bias in occupational AI-exposure measurement †Compare exposure estimates conditioned on which AI platforms workers use †Worker/platform usage + occupation data (pending) †Exposure measures shift with platform-selection assumptions †Exposure scores under selection models †D3: platform-selection bias in exposure measurement †Pending †Surfaces platform-selection bias standard exposure measures ignore †Working paper; authors and details pending †
18AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability-Automate red-teaming with multi-judge confirmation and transfer tests †Multi-judge confirmation pipeline; cross-model attack transfer (details pending) †Automated red-team prompts + multi-model targets (pending) †Multi-judge agreement + transferability findings (pending PDF) †Judge agreement; ASR / transfer rate †Red-teaming methodology; feeds #3’s judge-reliability design †Pending †Multi-judge confirmation plus transfer tests target the judge-unreliability weak link †Details pending PDF; preprint †
19Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting-Red-team T2I models with replay + semantic-preserving rewrites †In-context experience replay + semantic-preserving prompt rewriting †T2I prompt/attack suites (pending) †Improved T2I jailbreak/attack success via replay+rewrite (pending) †T2I ASR / safety filter bypass rate †Red-teaming methodology, T2I modality †Pending †Experience replay plus semantic-preserving rewrites as a reusable attack loop †T2I modality; details pending PDF †
20NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms (arXiv:2606.20408)H. Lee, D. Choi, B. Kim, H. Park, S.G. Kim, 2026Measure whether an adaptive adversary can drive a team of LLM operator agents to a physically unsafe state; harm grounded in an objective safety-function signal, not LLM-judged textClosed simulated nuclear-plant control room; five-role operator team; six critical safety functions; ≤10-turn adaptive sessions149 replayed sessions; four operator modelsAdaptive multi-turn attacks lose a CSF in 8.7-12.1% of sessionsASR_CSF; Wilson CIsAgentic methodology; rigor reference †Pending †Objective safety-function harm signal instead of LLM-judged text; Wilson CIs †Simulated control room; preprint †
21RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems (arXiv:2606.23927)-Unified adversarial eval across agentic architectures via graph NodeSpecStructure Identifier + probes45 agentic systems; 105 probesHigh SI accuracy; domain ASRs reportedAAR, ASR, utilityAgentic methodology †Pending †Unified graph NodeSpec abstraction across 45 agentic systems †Preprint; agentic scope only †
22Institutional Red-Teaming (arXiv:2607.07695)-Red-team deployment rules, not only modelsIABench-CA volunteer’s dilemma over five rules33,924 gamesRule changes move fatality 22-58 ppFatality; alignment gapRules as red-team target †Pending †Red-teams deployment rules, not models; rule changes move fatality 22-58 pp †Preprint; game-theoretic abstraction †
23Beyond Attack-Success Rate (arXiv:2607.07474)-Graded severity for tool-using agentsL0-L6 ordinal scale + judge reliabilityAgentDojoBinary ASR hides L4 residual harm under filtersASR; Cohen’s κEval design for #3 †Pending †L0-L6 graded severity exposes residual harm that binary ASR hides †Preprint; AgentDojo scope †
24EvalSafetyGap (arXiv:2606.30219)B.A. Uluırmak & R. Kurban, 2026Eval-safety gap survey + 10-model auditPRISMA + audit373 studiesGovernance-driven open/closed gapComposite safetyF-category related work †Pending †PRISMA-style survey of 373 studies plus a 10-model audit †Preprint; survey breadth over mechanism depth †
25Single-Layer RL Can Match Full-Parameter Training (arXiv:2607.01232)Z. Zhang et al., 2026Per-layer capacity for RLVR gainsSingle-layer GRPO; \(\mathcal{C}(k)\)Qwen-family; math/code/agentsMid-stack peak; often \(\mathcal{C}\geq 1\)\(\mathcal{C}(k)\); task avgsSpine C protocol; capability only †No repo logged; full detail on arXiv\(\mathcal{C}(k)\) protocol with fair-comparison LR rule; mid-stack peak across 7 models and 3 algorithms; rankings transfer across datasetsPreprint; Qwen family only; capability rewards, not refusal
26LoRA is All You Need for Safety Alignment of Reasoning LLMs (arXiv:2507.17075)Y. Xue, B. Mirzasoleiman, 2025Bypass safety tax via LoRA safety SFT; middle MLP up-proj bestRank-1 LoRA ablations (early/mid/late)Reasoning + safety benchesMiddle layers best tradeoffSafety vs reasoning metricsLocalization prior; under review ICLR 2026; no quant×lang †Not logged †Rank-1 LoRA sweep isolates middle MLP up-proj as the best safety/reasoning trade-off †Under review; no quantization or language axis †
27Alignment-Aware Quantization for LLM Safety (arXiv:2511.07842)Wee et al., 2025PTQ that preserves alignment via contrastive APC lossAAQ in PTQ pipeline; W4A4LLaMA/Qwen/Mistral; SafetyBench + AdvBenchClaims safety retained where PPL-only PTQ failsSafetyBench; ASRQuant×safety (English); preprint in venue check; no Multilingual MultiJail †Not logged †PTQ objective that optimizes alignment retention directly via contrastive APC loss †Preprint in venue check; English only †