SaFoLab : Security and Safe Foundation Model Systems
Pinned Loading
Repositories
- DRIFT Public
[NeurIPS 2025] The official implementation of the paper "DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents".
- Latent_Policy_Guard Public
Latent Policy Guard (LPG) — a guardrail model that performs semantic latent deliberation over dynamic safety policies. LPG compresses intent and risk reasoning into latent tokens and emits a compact policy-indexed verdict.
- ReasoningBomb Public
[CCS 2026] The official implementation of our CCS 2026 paper "ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models"
- MaskForge Public
- SafeVL Public
Official Repo for Paper: SafeVL: Driving Safety Evaluation via Meticulous Reasoning in Vision Language Models
- PW-OPSD Public
The official implementation of our preprint paper "When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning"
- Safety-Midtrain Public
Code for Safety Mid-Training: Internalizing LLM Safety as a Foundational Capability
- AgentDyn Public
The official implementation of the paper "AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?"
- ROM Public
The official implementation of our paper "ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention"
- A2ASecBench Public
Official code repository for "A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems" at ICLR 2026.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…