ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware
Paper • 2609.00955 • Published
PhD student at Beihang University — efficient LLM inference, speculative decoding, compute-in-memory.
custom_generate method for Transformers (5.19.x). Deterministic, bounded-regret acceptance policy for greedy assisted decoding; token-identical to strict verification at budget 0. Usage:model.generate(..., assistant_model=draft,
custom_generate="ynFeng/approximate-speculative-decoding", trust_remote_code=True,
assistant_asd_budget=2.0, assistant_asd_local_ratio=0.25, assistant_asd_max_mismatches=2)