news.volyx.in

Lossless LLM compression for efficient GPU inference via dynamic-length float (arxiv.org)

411 points by CharlesW · 463 days ago · 117 comments on HN

The AI summary for this story hasn't been generated yet — it's produced hourly. Check back soon. Meanwhile, read the discussion on HN.