kalomaze @kalomaze
[top, cut off] "...annoying here and it is making me want to kms" ๐ฌ1 โก2 ๐205
M @init_malachi ยท 7h
like per example or per batch
๐ฌ1 โก2 ๐268
kalomaze @kalomaze ยท 7h
per batch
it's not "A is compared to one B" but "A is compared to every B"
๐ฌ2 โก5 ๐249
M @init_malachi ยท 7h
interpreted it as contrastive learning
๐ฌ1 โก2 ๐142
kalomaze @kalomaze ยท 6h
i guess this is "contrastive classification" then?
๐ฌ1 โก5 ๐146
Ramesh Arvind @RameshArv1nd ยท 4h
Dumb question, if you're only using the contrastive loss how are you estimating CE loss (no head)? And also why abandon CE and not add the contrastive term as an aux loss.
I imagine for binary you could get away with some min/max sigmoidal diff across the batch
[cut off]
Note from Claude Sonnet 5
Continuation of the same ML training-technique thread as the prior screenshot (kalomaze discussing pairwise/contrastive classification loss formulation). Technical ML discussion, not AI-safety focused.
machine-learningtrainingcontrastive-learningloss-functionstechnical