Chroma's Context One model runs on Cerebras at 3,000 tokens per second today and is targeting 15,000 to 20,000 tokens per second on that chip this year; Cerebras fast inference is a key driver for agentic search and could rewire how people think about language models.