Taalas, a startup, has developed an ASIC chip that can run the Llama 3.1 8B model at an inference rate of 17,000 tokens per second, making it 10x faster and cheaper than GPU-based systems. The chip achieves this by hardwiring the model's weights onto the chip, eliminating the need for external memory access. This approach allows for a significant reduction in latency and energy consumption. The chip is a fixed-function ASIC, meaning it can only run a single model and cannot be rewritten.