Adding just 5 trainable parameters improve ImageNet-1K by +2.5 (+/- 0.5) percentage points - One Activation for One Network
A small ViT (~3M parameters), trained from scratch on ImageNet-1K:
• Model A: 3,048,232 params → 55.31% Top-1
• Model A+: 3,048,237 params → 57.49% Top-1
Only 5 extra trainable parameters, shared across the network, tested across 5 seeds, p-score 0.00056.
Everything else is identical: the same ViT architecture, training recipe, optimizer, and schedule. The improvement appears early, remains consistent across runs (≈ ±0.5 pp), across 100 epochs
A blog post on the findings.
I'd genuinely appreciate your thoughts and critiques.