Yuannuo Feng

PhD student at Beihang University — efficient LLM inference, speculative decoding, compute-in-memory.

Models

  • approximate-speculative-decoding — ASD as a custom_generate method for Transformers (5.19.x). Deterministic, bounded-regret acceptance policy for greedy assisted decoding; token-identical to strict verification at budget 0. Usage:
model.generate(..., assistant_model=draft,
    custom_generate="ynFeng/approximate-speculative-decoding", trust_remote_code=True,
    assistant_asd_budget=2.0, assistant_asd_local_ratio=0.25, assistant_asd_max_mismatches=2)

Papers

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for ynFeng/ynFeng