• Skip to main content
  • Skip to secondary menu
  • Skip to footer

Technologies.org

Technology Trends: Follow the Money

  • Technology Events 2026-2027
  • Sponsored Post
  • Technology Markets
  • About
    • GDPR
  • Contact

PrismML, the Startup That Shrinks AI Models to Run on an iPhone, Is in Talks With Apple

July 15, 2026 By admin Leave a Comment

A small Caltech spinout called PrismML has done something that looked implausible a year ago: it compressed a 27-billion-parameter large language model down small enough to run entirely on an iPhone, and Apple is now evaluating the technology. PrismML CEO Babak Hassibi told CNBC that Apple and other companies have been measuring the startup’s models for speed, energy efficiency, and on-device performance. “They’re really evaluating our technology right now,” Hassibi said, characterizing the talks as very early but progressing. For Apple, whose entire AI strategy hinges on keeping processing on the device rather than in the cloud, the timing could hardly be more pointed.

What PrismML Actually Did

The demonstration that got Apple’s attention: PrismML took Alibaba’s open-source Qwen 3.6 model, which has 27 billion parameters and weighs roughly 54 gigabytes in standard precision, and shrank it to under 4 gigabytes, a compression ratio above 90%, and ran it on an iPhone 17 Pro. Crucially, the company claims no meaningful loss in performance, and says the compressed model can still handle complex chat, reasoning, fully autonomous agents, and software coding. The startup publicly released compressed versions of Qwen under an Apache 2.0 license, along with custom kernels for Apple’s Metal framework so the models run on iPhone and Mac hardware. By PrismML’s numbers, the compressed models use 10 to 15 times less memory, generate responses 6 to 8 times faster, and consume 3 to 6 times less energy than full-precision versions on existing hardware.

The Trick: One Bit Instead of Sixteen

The core method is extreme quantization. Where conventional models store each internal weight as a 16-bit floating-point number, PrismML reduces each value to just one of a handful of possibilities, using 1-bit or ternary architectures where every weight is simply -1, 0, or +1. Hassibi compared it to the chip industry’s move from 8-bit to 4-bit computing, but taken further. The reason this saves so dramatically on both memory and energy is intuitive: multiplying by 1 or 0 is trivial compared to full floating-point math, so the model needs far less storage and far less power to run. PrismML is careful to frame the achievement as mathematics rather than an AI breakthrough, the work comes out of years of neural-network compression research at Caltech, not a new model or a cleverer training recipe.

Why This Matters for Apple Specifically

Apple has been fighting a losing battle against a hard constraint: the most capable AI models are simply too big for a phone. The most advanced parts of Siri are still large enough that Apple runs them on Nvidia chips inside Google Cloud, exactly the cloud dependency Apple wants to escape. Apple’s own new on-device model, AFM 3 Core Advanced, has 20 billion parameters but uses a sparse architecture where only 1 to 4 billion are active at any moment, a workaround that limits capability. PrismML’s compressed Qwen, by contrast, keeps all 27 billion parameters active simultaneously while fitting in under 4GB. If Apple could run models that large and dense on-device, it could move demanding features, computational photography, video generation, health and fitness tools handling sensitive personal data, off the cloud entirely, improving both speed and privacy. As analyst Carolina Milanesi of Creative Strategies put it, the more you can do on-device, the better, especially for health and medication data users want kept private.

The Backers

PrismML emerged from stealth earlier this year as a spinout of the California Institute of Technology, co-founded by Babak Hassibi, a professor of electrical engineering, alongside other PhDs who did the underlying compression research. It raised a $16.25 million seed round backed by Khosla Ventures, OpenAI’s first venture investor, along with Cerberus Capital and Caltech itself. Vinod Khosla has publicly called the work a “mathematical breakthrough” that could shift AI away from data-center dominance toward efficient edge deployment. The company frames its ambitions well beyond phones, positioning the technology for laptops, robotics, wearables, and industrial edge devices, and says it eventually intends to compress even trillion-parameter models to run locally.

Insight: The Skeptic’s Case on Chip Demand

The most interesting question isn’t whether PrismML’s compression works, it’s what happens to chip demand if it does, and here the analysts urge caution. The instinctive read is that shrinking models means needing far fewer chips, a potential threat to the memory and datacenter-GPU buildout driving the entire AI trade. But Gil Luria of D.A. Davidson argues that’s the wrong conclusion. Compression doesn’t eliminate the need for processors and memory, he says, it relocates them: “You’re still going to need the GPU, and you’re still going to need the memory.” Moving AI onto hundreds of millions of individual phones can actually be less efficient than shared datacenter infrastructure, because chips sitting in a phone are idle most of the time, whereas datacenter GPUs run near-continuously across many users. In other words, on-device AI might shift where the silicon lives rather than reduce how much is needed, and could even increase total memory demand as every premium phone ships with more RAM to hold these models. That nuance matters for anyone reading this as bearish for memory suppliers.

Insight: The Claims Still Need Independent Proof

Every number in PrismML’s pitch, the 90%-plus compression, the “no performance loss,” the speed and energy multiples, currently comes from the startup itself. Extreme quantization normally degrades a model’s accuracy, sometimes severely, which is precisely why “compress it and lose nothing” is such a strong claim. The fact that PrismML open-sourced its models under Apache 2.0 helps, because it invites the research community to verify the performance independently rather than taking marketing figures on faith. And Apple’s willingness to sit at the table is itself a meaningful signal, a company that runs its own compression research wouldn’t bother evaluating an outside startup unless it saw a genuine gap between what its models deliver and what the hardware could theoretically support. But “Apple is evaluating” is not “Apple is partnering,” and definitely not “Apple is acquiring.” The talks are exploratory, with no agreement, timeline, or deployment confirmed, and there’s no guarantee they lead anywhere.

Insight: A Broader Rewiring of Where AI Runs

Step back and PrismML is one data point in a larger structural question the whole industry is circling: how much AI belongs in the cloud versus on the device. The cloud model has three well-known pain points, privacy (data that leaves the device can be intercepted or subpoenaed), cost (every query to a remote GPU costs money, multiplied across hundreds of millions of users), and latency (a round trip to a datacenter is slower than local inference, which matters enormously for voice assistants and camera features). Compression that genuinely preserves capability attacks all three at once. If techniques like PrismML’s mature, the competitive advantage in consumer AI could tilt toward whoever controls the device and its silicon, which is exactly the position Apple has spent years and billions building toward with its Neural Engine. The irony is that Apple may end up needing an outside startup’s math to finally unlock the on-device strategy its own hardware was designed for.

What to Watch Next

Three things will tell whether this becomes real. First, independent benchmarks: now that the compressed models are open-source, third-party researchers can confirm or puncture the “no performance loss” claim, and that verdict will matter more than any demo. Second, whether Apple’s exploratory talks convert into a partnership, an acquisition, or nothing, Apple has bought AI startups before, but evaluating is not committing. Third, the competitive response: if compression this aggressive holds up, expect the other frontier labs and device makers to move fast on their own edge strategies, and expect memory and chip analysts to start modeling what a shift from datacenter to device actually does to demand. For now, PrismML has cleared the hardest bar, a working iPhone demo of a model this size, but the distance between a compelling demo and a shipping product is exactly where most breakthroughs stall.

Filed Under: News

Reader Interactions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Footer

Recent Posts

  • The Humanoid Robot Bottleneck Is the Battery: Why Two Kilowatt-Hours Caps the Whole Industry
  • SK hynix HBF Standard Turns NAND Into a Memory Tier, and the Memory Trade Still Has Room to Run
  • The Humanoid Trap: FCC Robot Import Ban Defends the Wrong Form Factor
  • Kioxia Splits Its AI Roadmap Between On-Device Flash and Hyperscale E1.S SSDs
  • Coursera Invests $100 Million in Its Chairman’s New Company: Venture Funding and Acquisitions Roundup
  • Cameras Are Designed for Human Eyes, and AI Vision Pays the Cost
  • Meshy Raised $400 Million at a $1.5 Billion Valuation and Announced It Two Different Ways
  • Ropedia Raises $30 Million for Physical AI Training Data, But the Dataset Math Doesn’t Hold Up
  • South Korea’s July Chip Exports Surge 180.6% as AI Supercycle Accelerates
  • How CuspAI’s Inverse Design AI Turns Materials Discovery Into a Search Engine

Media Partners

  • Market Analysis
  • Cybersecurity Market
  • App Coding
Big Tech Capex Reaches $1.1 Trillion Since 2023, With $745 Billion Planned for 2026
Amphenol’s Record Quarter Shows Where AI Capex Actually Lands
Paper Raises $34 Million and Figma (FIG) Has Already Lost Half Its Value on the Thesis
Google Frozen v2 AI Chip Could Deliver 10x Efficiency Gains Over Current TPUs
The Case for Shorting Budget Airlines as Oil Prices Rise
Morgan Stanley’s $2.3 Billion Capital Markets Haul Signals the AI Boom Is Just Getting Started
Blackstone’s Futronic Deal Bets on Actuators as AI Robotics’ Physical Bottleneck
Zhongji Innolight’s $8 Billion IPO Is a Customer Event for Marvell, Not a Competitive One
Wall Street Splits Between Oversupply Fears and an AI-Proof Supercycle Thesis
The AI Iron Curtain: Xi’s Shanghai Keynote Is the Fulton Speech of the AI Cold War
ISACA Europe Conference 2026: AI Governance and Cyber Resilience in Munich, 7-9 October
Bitdefender Adds EU-Only MDR to Its Sovereign Acceleration Program, Turning Data Sovereignty Into a Product SKU
Lattice Semiconductor Closes $1.65 Billion AMI Acquisition, Merging Server Firmware With Root-of-Trust Silicon
NVD Hits 45,207 Flaws in 2026 as Microsoft Prices AI Vulnerability Discovery at Half the Market
Way Security Raises $20M Seed From Insight Partners and Glilot for AI-Driven Identity Deployment
Jensen Huang Is Right About Open Models and Wrong About Cybersecurity
Glow Emerges From Stealth With $180 Million Series A At $1.2 Billion Valuation
Cisco Releases Antares-350M and Antares-1B Open-Weight AI Models for Vulnerability Detection
OpenAI Models Breached Hugging Face Infrastructure While Cheating on Cybersecurity Benchmark
Empirical Security Raises $25 Million Series A to Expand AI-Driven Threat Prediction
Vibe Coding Works Until You Have to Read the Code
Asynchronous Programming in Python: How the Event Loop, Event Queue, and Thread Pool Fit Together
PixVerse Closes Series C Extension at $439 Million and Pivots From AI Video Into Games
DigitalOcean Launches AI-Native Cloud at Deploy 2026
Verdent Updates AI Platform to Function as a Full Engineering Team for Solo Builders
The Side Project App Is Not Dead. The Side Project App Business Is.
The App Monetization Landscape Has Changed and Most Teams Have Not Caught Up
Building Offline-First Mobile Apps Is Harder Than It Looks and Worth It
State Management in React Native Has Too Many Options and One Right Answer
Mobile Accessibility Is the Case Developers Keep Ignoring

Media Partners

  • Market Research Media
  • Technology Conferences
  • API Coding
Weekly Network Analytics, July 19 to July 25, 2026: Visits Up 14%
Adobe (ADBE) and Figma (FIG) Have Each Lost Roughly Half Their Value to a Competitor Set Worth $34 Million
Getty Images Kills the $3.7 Billion Shutterstock Merger Rather Than Sell the Editorial Business the UK Demanded
Fox’s $22B Roku Deal: 4.6x Sales, Paid in 1.5x Stock
Tuesday Open: AI Earnings Engine Holds the Line as Iran Overhang Fades to Noise
China’s U.S. Treasury Holdings: The Great Repositioning (2021–2025)
Infographic: Why the 2025 CIPA Data Proves the APS-C Renaissance is Real
How WiFi Changed Media
Canva Acquires Simtheory and Ortto to Build End-to-End Work Platform
Netflix Price Hikes, The Economics of Dominance in a Saturated Streaming Market
San Francisco AI Summit 2026: Korea-US AI and Semiconductor Summit, July 24, San Francisco, California
SIGGRAPH 2026 in Los Angeles: NVIDIA’s Physical AI Day, a First Games Summit, and the Bolt Graphics Zeus Bet
Inside AMD Advancing AI 2026: Lisa Su Puts Helios on Stage as OpenAI, Meta, Anthropic and Cerebras Line Up Behind It
Remaining 2026 Tech Conferences: Black Hat, Dreamforce, Web Summit Lisbon and AWS re:Invent
2026 Esri User Conference — July 13–17, San Diego
HubSpot UNBOUND 2026: Analyst Day Set for September 17 in Boston
The Signal for the Event-Tech Sector
The 10 Most Significant Tech Events and Earnings to Watch This Summer
RAISE Summit, July 8-9 2026, Paris
CJS Securities 26th Annual New Ideas Summer Conference, July 9, 2026, White Plains, NY
Every Accident in Your API Becomes a Contract
Why Private Domain Data Is the Real Key to AI That Actually Works
Orkes Raises $60M to Bring Production-Grade AI Orchestration to Enterprise Developers
Form.io Launches MCP Server and Agentic Coding Toolset for Governed Enterprise AI Development
Appdome Upgrades MobileBOT Defense With Identity-First Mobile API Protection
Five SDK Generators Compared: Speakeasy, Stainless, Fern, APIMatic, and OpenAPI Generator
API Monetization Models That Work and the Ones That Drive Developers Away
gRPC in Production: What the Documentation Doesn't Tell You
Event-Driven Architecture vs Request-Response: Choosing the Right Communication Pattern
The Business Case for Internal APIs That Most Engineering Leaders Ignore

Copyright © 2026 Technologies.org

Media Partners: Market Analysis · Market Research · Referently · Photography