▲ 4 r/LLM

Adding just 5 trainable parameters improve ImageNet-1K by +2.5 (+/- 0.5) percentage points - One Activation for One Network

A small ViT (~3M parameters), trained from scratch on ImageNet-1K:

• Model A: 3,048,232 params → 55.31% Top-1
• Model A+: 3,048,237 params → 57.49% Top-1

Only 5 extra trainable parameters, shared across the network, tested across 5 seeds, p-score 0.00056.

Everything else is identical: the same ViT architecture, training recipe, optimizer, and schedule. The improvement appears early, remains consistent across runs (≈ ±0.5 pp), across 100 epochs

A blog post on the findings.

I'd genuinely appreciate your thoughts and critiques.

reddit.com
u/arun_ai — 6 days ago
▲ 4 r/ResearchML+1 crossposts

Can adding just 5 trainable parameters improve ImageNet-1K by +2.5 (+/- 0.5) percentage points?

A small ViT (~3M parameters), trained from scratch on ImageNet-1K:

• Model A: 3,048,232 params → 55.31% Top-1
• Model A+: 3,048,237 params → 57.49% Top-1

Only 5 extra trainable parameters, tested across 5 seeds, p-score 0.00056.

Everything else is identical: the same ViT architecture, training recipe, optimizer, and schedule. The improvement appears early (as shown in the figure in the attached drive link), remains consistent across runs (≈ ±0.5 pp)

Do you find this interesting?

If so, what would your hypothesis be?

A blog post on the findings

I'd genuinely appreciate your thoughts and critiques

drive.google.com
u/arun_ai — 8 days ago