Ternlight is a 7 MB embedding model that can run in a web browser, allowing for fast and offline semantic search without the need for API calls. The model uses ternary quantization-aware training and can embed text in milliseconds. It is available as a single npm package and can be used for various applications such as search, FAQ/intent matching, and clustering. The model's performance is comparable to other small embedding models, with a claimed embedding latency of around 5 ms.