← All topics

weight averaging

1 capture, most recent first.

Samuel Ratnam @eterecursion

— saved image

Samuel Ratnam @eterecursion · 29m
sidenote but averaging weights in ML works because useful circuits constructively interfere and task specific noise destructively interferes leading to better generalisation

[Quoted tweet]
Samuel Ratnam @eterecursion · 1h
this paper is so cool: arxiv.org/abs/2410.01463

you can get the benefits of data aggregation and personalisation without compromising on your privacy. You average everyone's A matrix to lear...

[Attached diagram, labeled 'FedSA-LoRA (Ours)': shows h = (arrows from) Pretrained weights W ∈ R^(d×d) [frozen, blue] plus B=0 [pink triangle] over A = N(0, σ²) [orange trapezoid, low-rank r] combined, with input x of dimension d at bottom. Legend: blue = frozen, orange = trainable - global, pink = trainable - local.]
Note from Claude Sonnet 5

Tweet by Samuel Ratnam explaining why weight averaging works in ML (constructive interference of useful circuits, destructive interference of task-specific noise), quoting his own earlier post about the FedSA-LoRA paper (arXiv:2410.01463) on federated personalization via averaging the LoRA A matrix while keeping B local, with the paper's architecture diagram attached.

machine learningfederated learningloraweight averagingtwitter