← All topics

llm inference

1 capture, most recent first.

@plugyawn

— saved image

Progyan @plugyawn · 12h
I spent the last six months trying to deconstruct Taalas's patents. We think it can be better: 100x fewer memfetches than their bitROM, better software by better quantization than their hardware team could.

So we wrote a compiler that:
> takes a huggingface checkpoint,
> quantizes the model
> descends the weights down metal layers to RTL and GDS, do your DRC, Yosys, PEX, make the electrical waveform execute a mat-vec from your huggingface checkpoint (thanks Cambricon tech papers)

in ~7000 lines of human-readable code.

we're looking for someone with contacts with a foundry/access to 7nm PDKs or contacts at a cheap EuroPTW shuttle? I'm broke and unemployed.

@itsclivetime this is what i think is the future of perplexity per picojoule. @zerohedge @zephyr_z9 you guys wanna see 40,000 tokens/sec?
Note from Claude Sonnet 5

Below the text are two embedded images of chip layout/routing diagrams (dense grids of colored horizontal and vertical traces in purple, orange, blue, red on dark background), presumably renders of the RTL/GDS output described in the tweet.

hardwarechipsasicllm inferencequantizationtwitter