Counterfactual evaluation for World Action Models
-
Updated
Aug 1, 2026 - Python
Counterfactual evaluation for World Action Models
Counterfactual dependability framework for testing whether language-model safety mechanisms are behaviorally load-bearing under removal, misrouting, bypass, and family-ablation controls.
Position-bias-aware ranking: estimating and correcting position bias to optimize Earnings Per Visitor (EPV) · simulation study · IPW & propensity modeling · Python
SAST-IR: evaluating LLM factual robustness with a stateful attacker and stateless target
Add a description, image, and links to the counterfactual-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the counterfactual-evaluation topic, visit your repo's landing page and select "manage topics."