The chip designer is banking on growing demand for inference, the data crunching that occurs when a user queries a chatbot, as it seeks to challenge Nvidia's dominance in the AI processor market.
Cerebras' flagship wafer-scale engine (WSE) is a single chip the size of a dinner plate containing trillions of transistors, a design that it says is more efficient than connecting thousands of smaller graphics processors together, as Nvidia does.
By placing memory directly on the chip, the WSE is built to accelerate inference and reduce the data-transfer delays associated with conventional graphics processors that rely on separate high-bandwidth memory.
Placing memory directly on the chip has lessened the impact of surging memory prices and placed it in a better position to compete with Nvidia, Cerebras CEO Andrew Feldman told Reuters in an interview.
"Nvidia's prices have gone through the roof because of HBM prices," Feldman said, referring to the high-bandwidth memory included with AI processors. "This is a battleground, and if they can't deliver or they're having significant component price increases, of course that helps."