Skip to content

Add PUMA (Preprint'26) + create 2026 section under Reasoning in LLMs - #75

Open
ZhishanQ wants to merge 1 commit into
atfortes:mainfrom
ZhishanQ:add-puma
Open

Add PUMA (Preprint'26) + create 2026 section under Reasoning in LLMs#75
ZhishanQ wants to merge 1 commit into
atfortes:mainfrom
ZhishanQ:add-puma

Conversation

@ZhishanQ

Copy link
Copy Markdown

Two small changes to 馃敜 Reasoning in Large Language Models:

  1. Create new `### 2026` year section at the top (above the existing `### 2025`), matching the year-ordered structure used throughout the repo.
  2. Add PUMA as the inaugural 2026 entry, following the existing format (**[Title.](url)** [[code](url)] then *Authors.* Venue'YY).

PUMA (Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models) is a plug-and-play inference-time early-exit framework for Large Reasoning Models that uses reasoning-level semantic redundancy as a complementary stopping signal to answer-level confidence/consistency. Architecture: lightweight Redundancy Detector (fine-tuned Qwen3-Embedding-0.6B) + Answer Verification window + Loop Breaker fallback.

Results across 5 LRMs (DeepSeek-R1-Distill-Qwen-7B/14B/32B, Llama-3.1-Nemotron-Nano-8B, Qwen3-30B-A3B-Thinking) and 5 reasoning benchmarks (MATH-500, AIME24, AIME25, OlympiadBench, GPQA-Diamond): 26.2% average token reduction with preserved (slightly improved) accuracy; 1.40脳 / 1.28脳 wall-clock speedup on DS-7B/14B. Also generalizes to LiveCodeBench (code) and zero-shot VLM reasoning (MathVista, MathVision).

Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant