The production engine for directional ablation. Unalign / remove models censorship efficiently on any hardware.
-
Updated
Jun 28, 2026 - Python
The production engine for directional ablation. Unalign / remove models censorship efficiently on any hardware.
Educational analysis of LLM alignment, safety behavior, and framing-sensitive response patterns.
SoftPrompt-IR is a low-level symbolic annotation layer for LLM prompts, making intent strength, direction, and priority explicit. It is not a DSL or framework, but a minimal, composable way to reduce ambiguity, improve safety, and structure prompts.
DSPy framework for detecting and preventing safety override cascades in LLM systems. Research-grade implementation for studying when completion urgency overrides safety constraints.
Adversarial evaluation framework for embodied and agentic AI — failure-first methodology, jailbreak corpus, VLA red-teaming, and policy research.
🌐 Detect and prevent safety overrides in LLM systems with this DSPy-based framework, ensuring actions align with safety constraints.
Research on multi-turn conversational manipulation of LLMs — can a small specialist attacker beat scale? Defensive AI-safety red-team methodology with programmatic judges. Pre-alpha.
Contract-enforced sandbox for studying AI agent self-replication safety
2x2 fine-tuning datasets (RU/EN x toxic/polite) for LLM safety-degradation research
Explore glider aviation safety through in-depth data analysis. This project leverages incident reports and manufacturing data, utilizing Python and Jupyter Notebooks for trend identification, risk assessment, and safety enhancement in glider aviation.
Add a description, image, and links to the safety-research topic page so that developers can more easily learn about it.
To associate your repository with the safety-research topic, visit your repo's landing page and select "manage topics."