• Skip to main content
  • Skip to secondary menu
  • Skip to footer

Technologies.org

Technology Trends: Follow the Money

  • Technology Events 2026-2027
  • Sponsored Post
  • Technology Markets
  • About
    • GDPR
  • Contact

PrismML, the Startup That Shrinks AI Models to Run on an iPhone, Is in Talks With Apple

July 15, 2026 By admin

A small Caltech spinout called PrismML has done something that looked implausible a year ago: it compressed a 27-billion-parameter large language model down small enough to run entirely on an iPhone, and Apple is now evaluating the technology. PrismML CEO Babak Hassibi told CNBC that Apple and other companies have been measuring the startup’s models for speed, energy efficiency, and on-device performance. “They’re really evaluating our technology right now,” Hassibi said, characterizing the talks as very early but progressing. For Apple, whose entire AI strategy hinges on keeping processing on the device rather than in the cloud, the timing could hardly be more pointed.

What PrismML Actually Did

The demonstration that got Apple’s attention: PrismML took Alibaba’s open-source Qwen 3.6 model, which has 27 billion parameters and weighs roughly 54 gigabytes in standard precision, and shrank it to under 4 gigabytes, a compression ratio above 90%, and ran it on an iPhone 17 Pro. Crucially, the company claims no meaningful loss in performance, and says the compressed model can still handle complex chat, reasoning, fully autonomous agents, and software coding. The startup publicly released compressed versions of Qwen under an Apache 2.0 license, along with custom kernels for Apple’s Metal framework so the models run on iPhone and Mac hardware. By PrismML’s numbers, the compressed models use 10 to 15 times less memory, generate responses 6 to 8 times faster, and consume 3 to 6 times less energy than full-precision versions on existing hardware.

The Trick: One Bit Instead of Sixteen

The core method is extreme quantization. Where conventional models store each internal weight as a 16-bit floating-point number, PrismML reduces each value to just one of a handful of possibilities, using 1-bit or ternary architectures where every weight is simply -1, 0, or +1. Hassibi compared it to the chip industry’s move from 8-bit to 4-bit computing, but taken further. The reason this saves so dramatically on both memory and energy is intuitive: multiplying by 1 or 0 is trivial compared to full floating-point math, so the model needs far less storage and far less power to run. PrismML is careful to frame the achievement as mathematics rather than an AI breakthrough, the work comes out of years of neural-network compression research at Caltech, not a new model or a cleverer training recipe.

Why This Matters for Apple Specifically

Apple has been fighting a losing battle against a hard constraint: the most capable AI models are simply too big for a phone. The most advanced parts of Siri are still large enough that Apple runs them on Nvidia chips inside Google Cloud, exactly the cloud dependency Apple wants to escape. Apple’s own new on-device model, AFM 3 Core Advanced, has 20 billion parameters but uses a sparse architecture where only 1 to 4 billion are active at any moment, a workaround that limits capability. PrismML’s compressed Qwen, by contrast, keeps all 27 billion parameters active simultaneously while fitting in under 4GB. If Apple could run models that large and dense on-device, it could move demanding features, computational photography, video generation, health and fitness tools handling sensitive personal data, off the cloud entirely, improving both speed and privacy. As analyst Carolina Milanesi of Creative Strategies put it, the more you can do on-device, the better, especially for health and medication data users want kept private.

The Backers

PrismML emerged from stealth earlier this year as a spinout of the California Institute of Technology, co-founded by Babak Hassibi, a professor of electrical engineering, alongside other PhDs who did the underlying compression research. It raised a $16.25 million seed round backed by Khosla Ventures, OpenAI’s first venture investor, along with Cerberus Capital and Caltech itself. Vinod Khosla has publicly called the work a “mathematical breakthrough” that could shift AI away from data-center dominance toward efficient edge deployment. The company frames its ambitions well beyond phones, positioning the technology for laptops, robotics, wearables, and industrial edge devices, and says it eventually intends to compress even trillion-parameter models to run locally.

Insight: The Skeptic’s Case on Chip Demand

The most interesting question isn’t whether PrismML’s compression works, it’s what happens to chip demand if it does, and here the analysts urge caution. The instinctive read is that shrinking models means needing far fewer chips, a potential threat to the memory and datacenter-GPU buildout driving the entire AI trade. But Gil Luria of D.A. Davidson argues that’s the wrong conclusion. Compression doesn’t eliminate the need for processors and memory, he says, it relocates them: “You’re still going to need the GPU, and you’re still going to need the memory.” Moving AI onto hundreds of millions of individual phones can actually be less efficient than shared datacenter infrastructure, because chips sitting in a phone are idle most of the time, whereas datacenter GPUs run near-continuously across many users. In other words, on-device AI might shift where the silicon lives rather than reduce how much is needed, and could even increase total memory demand as every premium phone ships with more RAM to hold these models. That nuance matters for anyone reading this as bearish for memory suppliers.

Insight: The Claims Still Need Independent Proof

Every number in PrismML’s pitch, the 90%-plus compression, the “no performance loss,” the speed and energy multiples, currently comes from the startup itself. Extreme quantization normally degrades a model’s accuracy, sometimes severely, which is precisely why “compress it and lose nothing” is such a strong claim. The fact that PrismML open-sourced its models under Apache 2.0 helps, because it invites the research community to verify the performance independently rather than taking marketing figures on faith. And Apple’s willingness to sit at the table is itself a meaningful signal, a company that runs its own compression research wouldn’t bother evaluating an outside startup unless it saw a genuine gap between what its models deliver and what the hardware could theoretically support. But “Apple is evaluating” is not “Apple is partnering,” and definitely not “Apple is acquiring.” The talks are exploratory, with no agreement, timeline, or deployment confirmed, and there’s no guarantee they lead anywhere.

Insight: A Broader Rewiring of Where AI Runs

Step back and PrismML is one data point in a larger structural question the whole industry is circling: how much AI belongs in the cloud versus on the device. The cloud model has three well-known pain points, privacy (data that leaves the device can be intercepted or subpoenaed), cost (every query to a remote GPU costs money, multiplied across hundreds of millions of users), and latency (a round trip to a datacenter is slower than local inference, which matters enormously for voice assistants and camera features). Compression that genuinely preserves capability attacks all three at once. If techniques like PrismML’s mature, the competitive advantage in consumer AI could tilt toward whoever controls the device and its silicon, which is exactly the position Apple has spent years and billions building toward with its Neural Engine. The irony is that Apple may end up needing an outside startup’s math to finally unlock the on-device strategy its own hardware was designed for.

What to Watch Next

Three things will tell whether this becomes real. First, independent benchmarks: now that the compressed models are open-source, third-party researchers can confirm or puncture the “no performance loss” claim, and that verdict will matter more than any demo. Second, whether Apple’s exploratory talks convert into a partnership, an acquisition, or nothing, Apple has bought AI startups before, but evaluating is not committing. Third, the competitive response: if compression this aggressive holds up, expect the other frontier labs and device makers to move fast on their own edge strategies, and expect memory and chip analysts to start modeling what a shift from datacenter to device actually does to demand. For now, PrismML has cleared the hardest bar, a working iPhone demo of a model this size, but the distance between a compelling demo and a shipping product is exactly where most breakthroughs stall.

Filed Under: News

Footer

Recent Posts

  • Top 10 Emerging Technologies in 2026
  • The World Economic Forum and Forrester Can’t Agree on What Counts as Emerging Technology in 2026
  • Snap’s AI Glasses, Faraday Future’s Robot Push, and Fresh AI Funding Lead the Sept. 16-17 Tech Wire
  • Bending Spoons Buys Miro at a 90% Discount
  • Morning Tech Digest, September 10, 2026: Chinese AI Chip Prices Up 20% to 50% on HBM Costs, Nasdaq’s $100 Million Kraken Bet
  • Apple Watch Series 12 and Ultra 4: The Hard Part of Audio Intelligence Is Everyone Not Wearing the Watch
  • Apple iPhone 18 Pro: The Base Price Rose $100, the Top Storage Step Rose $600
  • Anthropic Walks Away From $6 Billion Decart Acquisition
  • Collapse of Kenya’s academic ghostwriting industry
  • OpenAI Chief Scientist Jakub Pachocki Says No Lab Has Solved Alignment Well Enough to Scale at Full Speed

Media Partners

  • Market Analysis
  • Cybersecurity Market
  • App Coding
Semiconductor Revenue Hits Record $425B in Q2 2026, but Omdia’s $500B Q3 Forecast Implies Growth Halves
AI Extinction Warnings Went Global in Six Days. Nothing in the Technology Changed.
Anthropic Walks Away From $6 Billion Decart Acquisition: The Deal Was About Inference Cost, Not World Models
VR Status Report 2026: Quest Sales Keep Falling While Smart Glasses Take the Money
The Case That the US Can Grow Out of $40 Trillion in Debt: Three Conditions the Clinton Surpluses Actually Met
The $40 Trillion Debt: Why AI Capex Raises Treasury Borrowing Costs Faster Than It Raises the Tax Base
Rockefeller Center Has Been a Credit Instrument for Forty Years: From the 1985 REIT to the $3.5 Billion 2024 CMBS
Who Insures the AI Buildout? $30 Billion Campuses Meet a $3.5 Billion Ceiling
Retail Earnings Week: The 1.65% Real Sales Number Behind the 5% Headline
SanDisk and Marvell Top Our Hot Stocks List: Two-Thirds of FY2028 NAND Bits Are Already Contracted
Nvidia’s Huang Calls Cybersecurity AI’s Next Market, the One Demand Source AI Creates for Itself
OpenAI Agents Beat a GET-Only Sandbox Using a 25-Year-Old Wiki and a Fake Azure Hostname
Billington CyberSecurity Summit 2026: AI-Enabled Threats Take Center Stage in Washington, Sept. 8-10
Cybersecurity Stocks Rally: The 122-Point Spread Between Fortinet and Zscaler Says This Is Not a Sector Trade
CrowdStrike Fal.Con 2026: 150+ Sponsors and 10,000 Attendees at Mandalay Bay, August 31 – September 3
Datavault AI Will Pay $94.5 Million in Cash for CyberCatch, a Company With Roughly $230,000 in Annual Revenue
Oligo Security Raises $60 Million as Runtime Vendors Turn Post-Mythos Into a Market Category
ISACA Europe Conference 2026: AI Governance and Cyber Resilience in Munich, 7-9 October
Bitdefender Adds EU-Only MDR to Its Sovereign Acceleration Program, Turning Data Sovereignty Into a Product SKU
Lattice Semiconductor Closes $1.65 Billion AMI Acquisition, Merging Server Firmware With Root-of-Trust Silicon
Application Performance Optimization: Where Most Teams Waste Their Time
AI App Builders by Use Case: Lovable, Bolt.new, Replit Agent, Softr, FlutterFlow and v0
AI App Builders Reviewed: Lovable, Base44, Bolt, Replit and v0 Compared
Cloudflare Kitesurf: An Agent-First Browser That Uses 3-7x Less Memory Than Chromium
Vibe Coding Works Until You Have to Read the Code
Asynchronous Programming in Python: How the Event Loop, Event Queue, and Thread Pool Fit Together
PixVerse Closes Series C Extension at $439 Million and Pivots From AI Video Into Games
DigitalOcean Launches AI-Native Cloud at Deploy 2026
Verdent Updates AI Platform to Function as a Full Engineering Team for Solo Builders
The Side Project App Is Not Dead. The Side Project App Business Is.

Media Partners

  • Market Research Media
  • Technology Conferences
  • API Coding
The Economist Is Right About a Million AI Jobs. It’s a Construction Boom, Not a Tech Boom.
AI Slop Earns Higher CPMs Than Clean Inventory: Why the Ad Market Cannot Fix the Web It Funds
Weekly Network Analytics, July 19 to July 25, 2026: Visits Up 14%
Adobe (ADBE) and Figma (FIG) Have Each Lost Roughly Half Their Value to a Competitor Set Worth $34 Million
Getty Images Kills the $3.7 Billion Shutterstock Merger Rather Than Sell the Editorial Business the UK Demanded
Fox’s $22B Roku Deal: 4.6x Sales, Paid in 1.5x Stock
Tuesday Open: AI Earnings Engine Holds the Line as Iran Overhang Fades to Noise
China’s U.S. Treasury Holdings: The Great Repositioning (2021–2025)
Infographic: Why the 2025 CIPA Data Proves the APS-C Renaissance is Real
How WiFi Changed Media
AGNTCon + MCPCon Europe Opens in Amsterdam September 17-18 With Stateless MCP on the Keynote Stage
AGNTCon + MCPCon North America 2026: Agentic AI Foundation Flagship Lands in San Jose, October 22-23
Cloudflare Connect 2026: Full Agenda, $595 Pass and 100+ Sessions at Moscone West, October 19-21
September 2026 Investor Conference Calendar
FPGAworld Conference 2026: Stockholm, 8 September
swampUP 2026: JFrog’s Software Supply Chain Conference Hits The Glasshouse in New York, September 1-3
Node.js Interactive 2026, August 12–13, 2026, Atlanta, Georgia
Q4 2026 Semiconductor and Memory Conferences: Dates, Locations, Who Presents
FMS 2026 in Santa Clara: Kioxia, Samsung, SanDisk and SK Hynix Offer Four Incompatible Fixes for the AI Memory Wall
San Francisco AI Summit 2026: Korea-US AI and Semiconductor Summit, July 24, San Francisco, California
API Monetization Models: How Companies Actually Charge for Access
API Testing Strategies: What to Test and When
AI Platforms for Designing APIs in 2026: Spec Editors, SDK Generators, MCP Builders and AI Gateways Reviewed
Every Accident in Your API Becomes a Contract
Why Private Domain Data Is the Real Key to AI That Actually Works
Orkes Raises $60M to Bring Production-Grade AI Orchestration to Enterprise Developers
Form.io Launches MCP Server and Agentic Coding Toolset for Governed Enterprise AI Development
Appdome Upgrades MobileBOT Defense With Identity-First Mobile API Protection
Five SDK Generators Compared: Speakeasy, Stainless, Fern, APIMatic, and OpenAPI Generator
API Monetization Models That Work and the Ones That Drive Developers Away

Copyright © 2026 Technologies.org

Media Partners: Market Analysis · Market Research · Referently · Photography