What It Is
A non-autoregressive text classifier that fine-tunes Qwen2.5-0.5B with LoRA to resolve 15+ task families in a single forward pass — trading generative decoding for fast, calibrated label prediction.
How It Works
A LoRA adapter is trained on top of the Qwen2.5-0.5B backbone. Instead of decoding tokens one at a time, the pooled representation feeds a unified classification head that emits task-family logits in one pass, then calibrates them for reliable confidence.
Key Highlights
- Qwen2.5-0.5B backbone adapted with LoRA for parameter-efficient fine-tuning
- Non-autoregressive single-pass inference across 15+ task families
- 8.9x batch throughput speedup versus generative decoding baselines
- Strong calibration with a measured 0.030 test ECE
- Unified classification head replacing per-task bespoke pipelines
Why It Matters
Shows how to turn a generative LLM into a fast, calibrated classifier — the kind of efficiency work that makes models cheap enough to ship.
Stack