Apple is evaluating technology from PrismML that could shrink powerful AI models to run directly on iPhones, the startup's CEO told CNBC.
The Khosla Ventures-backed company, spun out from Caltech, released compressed versions of Alibaba's open-source Qwen model on Tuesday. PrismML reduced the model from roughly 54GB to less than 4GB, allowing all 27 billion parameters to run on an iPhone 15 or newer.
CEO Babak Hassibi said Apple and other companies are measuring the startup's models for speed, energy efficiency and performance on devices. "They're really evaluating our technology right now," he said of Apple, characterizing the discussions as very early but "progressing nicely."
The compression breakthrough
PrismML shrinks AI models by drastically simplifying how their internal information is stored — reducing each value from 16 bits to just one or three possible values. This cuts memory requirements while generating responses six to eight times faster and consuming three to six times less energy.
The compressed models lose a few percentage points of overall performance, with factual recall weakening before skills like reasoning and coding, Hassibi acknowledged.
The technology emerged from Hassibi's research group at Caltech, which owns the underlying patents and licenses them exclusively to PrismML. The company raised a $16.25 million seed round in March.
Apple's on-device strategy
The breakthrough could help Apple make Siri faster and more private by keeping AI processing on-device instead of sending requests to the cloud. The release comes one day after Apple opened the public beta of iOS 27, giving iPhone owners broad access to the company's Siri overhaul.
Apple already runs parts of its AI system locally, including translation and summarization. More complex requests are routed to Apple's private cloud infrastructure or outside models.
Carolina Milanesi from Creative Strategies said smaller models could let Apple move demanding features like computational photography and health tools onto the iPhone. "The more you can do on device, the better it is," she said.
PrismML plans to compress Google's Gemma model next, followed by much larger models from frontier labs that typically require datacenter hardware.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.