# Zach Rattner's Marketing & Speaker Knowledge Base Zach Rattner is the "AI Founder for the Physical World." He is a battle-tested CTO, inventor holding 31 granted US patents, and top AI keynote speaker for insurance, logistics, and enterprise operations. He delivers high-impact, technical keynotes bridging framework innovation with measurable business ROI. Last updated: 2026-09-06T14:57:57.199Z --- ## Page: Top AI innovation keynote speaker Zach Rattner **URL:** / **Description:** Zach Rattner transforms organizations with actionable AI blueprints. A seasoned CTO and inventor, he delivers keynotes that turn AI into measurable results. ## AI founder for the physical world # Zach Rattner Invite Zach to Speak See Zach in Action ## AI Founder | Author | Speaker | Leader AI should be empowering, not overwhelming. I help organizations around the world apply AI to get real results in the physical world. ## About Zach Zach Rattner's keynotes give leaders a working blueprint for turning AI into results across physical industries like insurance, logistics, and home services. As the Chief Technology Officer and Co-Founder of Yembo, an AI-powered virtual inspection platform used in 20+ countries, Zach's work has helped organizations complete 3x more inspections per day while improving accuracy and the customer experience. Zach holds 31 granted US patents and has guided teams worldwide to adopt AI responsibly and effectively. Zach leads innovation with a compassionate approach, because technology only matters when it makes lives better. ### Book Zach to speak Audiences walk away knowing exactly how to use AI to work smarter, innovate faster, and create lasting impact. Let's Talk About Your Event AI deployed in 1+ Countries Delivering AI deployments internationally, bringing practical expertise to organizations worldwide. Over 5k Professionals trained worldwide I have trained teams across the globe to apply AI tools that save them time, cut their costs, and widen what they can take on. Currently 1 Granted US patents with several more pending Pioneering the invention and patents of technology used in home services, insurance, data security, finance and the future of work. ## Join 5,000+ professionals who've learned from me DISCOVER KEYNOTES THAT DRIVE REAL CHANGE ## Sectors ### Construction and remodeling Seven in ten AEC leaders expect AI to change this industry, and about one in ten has shipped it. That gap is a data problem, and it starts on the job site. I show your teams how to turn phone video of a site into measurements and scope you can actually bid from. Explore AI for Construction and remodeling 1 ### Home services You cannot hire your way out of a 110,000 technician shortage, but you can stop spending the technicians you have on paperwork. I show your leaders which parts of the day AI can take off a crew and which parts your customers are actually paying for. Explore AI for Home services 2 ### Insurance Compliance is the only thing that will let you ship your AI program. I have put computer vision into production claims and inspection workflows, and I show your teams where automation cuts cycle time and where it quietly costs more than it saves. Explore AI for Insurance 3 ### Data security Your AI ban made AI use invisible, and invisible is the expensive kind. I have taken an AI platform through ISO 27001, SOC 2, GDPR, and NIST 800-171, so I can show your leaders what genuinely raises risk and how to ship without waiting for the audit. Explore AI for Data security 4 ### Customer service Deflection counts the customers you got rid of. Your customers can tell the difference, and so can your CSAT. I show contact center leaders where AI genuinely resolves a problem, where it only hides one, and how to hand agents the routine without losing the judgment. Explore AI for Customer service 5 ### Interior design Your taste is safe. The hours you bill around it are the problem, and nobody is talking about that honestly. I show your designers what a room walkthrough can become and what that means for how you charge. Explore AI for Interior design 6 ### Moving and relocation Your close rate is your growth problem, and your estimators are the ones who can fix it. I built the video survey AI that movers in 20+ countries run every day, and I show your teams what it actually changed about surveying, estimating, and dispatch. Explore AI for Moving and relocation 7 ### Fine art logistics The job here is having the record you wish you had taken, before the object moved. I show registrars and shippers how computer vision turns video into a documented, item-level inventory without slowing down the crew that handles the work. Explore AI for Fine art logistics 8 #### Phil Hunter Entrepreneur #### Sue Hurrell Specialist, Events Planning at AICPA & CIMA #### Prachi Joshi Program Manager, Project Management Institute Rattner closed with a truth many avoid: no one has all the answers yet — our opportunity lies in curiosity. Luke Baldwin UI / UX Designer at People's Partnership Zach's insights were practical, engaging, and very well received by the attendees. FIDI 39 Club A great talk that we are still talking about! Southwest Movers Association Really enjoyed your talk this morning. It was easily one of the most useful sessions I've seen at Brighton. Katie Liguori Founder, KALVA Creative You made me more hopeful and inspired me to be more playful and experimental with AI. Julie Sun Lead UX Consultant, Softwire ### Podcast appearances VIEW PODCASTS ### Bring Zach to your next event BOOK ZACH ## Products for your AI journey ### AI Conversation Starter The AI Conversation Starter cards are designed to help leaders, teams, and innovators work out what AI actually changes about their business. BUY NOW ### Grow Up Fast: Lessons from an AI Startup The untold narrative of our era is that there are yet untraveled paths to tread and new discoveries to uncover in the world of AI. BUY NOW ## Yembo is trusted around the world every day I am CTO and Co-Founder at Yembo, providing AI-driven property inspections for moving and insurance companies worldwide. Each dot on the globe represents a Yembo survey performed in a typical month. ## Podcast appearances I talk about AI, leadership, and innovation across podcasts in tech, business, and physical industries. SEE ALL August 2026 ### Yembo completes CMMC 2.0 Level 2 self-assessment, keeping military moves fully supported July 2026 ### Zach featured as Intruder pioneers AI-powered penetration testing June 2026 ### Zach featured in Meetings Today on keeping sensitive data secure with AI June 2026 ### Zach on keeping contract data secure while using AI, Smart Meetings March 2026 ### Zach joins Silicon Valley Speakers to keynote on AI and innovation February 2026 ### Zach's CSUF collaboration uses AI to help visitors get personal with art February 2026 ### Graebel and Yembo launch industry-first AI technology to transform moving claims December 2025 ### Zach and Rita Sus unveil Lightwall, an AI installation that reacts to you December 2025 ### Zach named an AI instructor at AICPA & CIMA's Business Learning Institute October 2025 ### SYCN Auto Logistics and Yembo announce partnership to streamline relocation services September 2025 ### A marriage of love - and logistics, FIDI Focus May 2025 ### Zach joins CSU Fullerton's Applied AI Advisory Board April 2025 ### Zach on why AI concerns run deeper than technology replacing people, FM Magazine December 2024 ### Zach releases Scourhead, a free open-source AI agent that streamlines research August 2024 ### Virginia Tech spotlights Zach, the alumnus reshaping moving with AI May 2024 ### Yembo partners with Atlas Van Lines, San Diego Business Journal May 2024 ### Zach asks what happens after AI enters the workforce, IAM Portal March 2024 ### AI-narrated books are here. Are humans out of a job? San Diego Union Tribune March 2024 ### IEEE Spectrum profiles Zach and the AI tool making moving day easier NEWER OLDER --- ## Page: AI keynote speaker for business events | Zach Rattner **URL:** /speaking **Description:** Book Zach Rattner, CTO and patented inventor, to speak on AI strategy, adoption, and automation. Keynotes and workshops for events, associations, and teams. # AI keynote speaker I am a working CTO with 31 granted patents, and I have shipped AI in production across 20+ countries. I speak from what I have actually built, so your teams leave with the parts that worked, the parts that did not, and a way to tell those apart before your budget finds out. Invite Zach to Speak See Zach in Action ## Inspire action. Apply AI. Deliver results. Create lasting impact. ### AI should make people more powerful, not replace them. I still build this for a living, which means the framework I give you has been tested against real production systems rather than a slide. Your people leave able to apply it in physical industries, where AI has to work alongside trucks, claims, and customers rather than in a demo. Book a Results-Driven Keynote ## Speaking programs 01 Why AI fails (and how to fix it) Your company has AI projects running right now, and you probably cannot say which of them will pay for itself. I have shipped AI in production across 20+ countries and watched which initiatives return money and which quietly become someone's side project. In this keynote I show your teams how to tell those apart before the budget is spent, and how to move the survivors into daily operations. ### Who it's for: Companies with AI on the roadmap and no agreed way to judge it. Strongest fit for insurance, home services, logistics, financial services, and real estate, where the work is physical and the AI has to survive contact with it. ### Key learning outcomes: - Leave knowing which three of your AI projects are worth funding - Judge an AI initiative before it has spent your budget - See how teams across 20+ countries moved AI from pilot to production 02 Thriving through change Your people are watching you to find out whether AI is a threat or a tool, and they are reading your hesitation as an answer. I have led an engineering organization through exactly this, badly at first. In this keynote I give you the frameworks I use to make decisions while the ground is still moving, and to keep a team steady enough to act on them. ### Who it's for: Executives, managers, and team leads whose organizations are adopting AI faster than they can explain it. Built for the moment when your people need a decision and you do not yet have certainty. ### Key learning outcomes: - Make the call before you have all the information, and defend it - Keep your team willing to try things that might not work - Judge a new AI tool without waiting for the market to decide 03 Automate the ordinary, elevate the extraordinary Your best people are spending their week on work that a machine should be doing, and everyone in the room knows which work it is. I bring real examples from building Yembo, including the ones that did not go how I expected, and show your teams how to hand off the routine without handing off the judgment. AI should make your people more powerful, not smaller. ### Who it's for: All-hands and cross-functional audiences, especially operations, customer service, HR, and marketing. Right when you want the whole room to leave with something they can try on Monday rather than a strategy they have to wait for. ### Key learning outcomes: - Name the five tasks in your week worth automating first - Start a pilot small enough that failing costs you nothing - Give your team a reason to reach for AI instead of bracing for it 04 Add a working session AI Strategy Roundtable - One hour with your senior leaders, working on your actual opportunities rather than generic ones. I prioritize the list with you and leave the room aligned on what happens next. AI Action Plan - Two hours in which your teams build a 30-day roadmap for their first pilots. Everyone leaves with a named owner, a number to hit, and a task for Monday. AI Opportunity Audit - Four hours pulling apart the workflows you run today to find what is worth automating. You keep a written opportunity map with the savings and the payback math laid out. Momentum Session - A virtual session 60 days on, because this is where most initiatives quietly stall. We go through what shipped, clear whatever is blocking the rest, and reset the next set of goals. ## Featured engagements and impact ### How insurance carriers can leverage AI #### The Challenge Analog claims processes are actively losing customers who now expect instant, digital-first service. #### The Solution Streamline the claims lifecycle by replacing subjective guesswork with objective, AI-verified visual data. "Zach Rattner was fantastic to work with from start to finish. His presentation was very engaging and left the audience buzzing." Sofie Abramowitz, Key Media ### Turning AI anxiety into your creative advantage #### The Challenge Designers, technologists, and product leaders are currently paralyzed by the rapid evolution of AI. #### The Solution This talk completely reframed the AI learning curve from a threat into a competitive advantage. "You made me more hopeful and inspired me to be more playful and experimental with AI." - Julie Sun, Softwire ### Maximizing economic return on AI #### The Challenge Businesses are pouring money into AI experiments without a clear line from adoption to profit. #### The Solution A practical keynote showing exactly where AI pays off, from high-ROI workflows to measurable savings. "A great talk that we are still talking about!" Southwest Movers Association ### 10 Reasons to be optimistic about AI #### The Challenge Guiding executive leadership to confidently adopt practical AI use-cases despite uncertainty. #### The Solution Inspirational keynote breaking down complex AI shifts into approachable human-centered tactics. "Zach was awesome - the highlight." Attendee, AICPA & CIMA ENGAGE '24 ### Turning AI pilots into daily operations #### The Challenge Global mobility leaders had AI tools in hand and still finished every day with the same backlog. #### The Solution A working session in Osaka that mapped AI onto the workflows their businesses actually run, closing with a 30-day rollout plan each attendee could start on Monday. "Zach's insights were practical, engaging, and very well received by the attendees." FIDI 39 Club ## Let me give your audience something they can use on Monday Book a Results-Driven Keynote ## Podcast Appearances ### Boss It Mark Edwards ### Breakthrough Innovation JL Heather ### Business Bros Hernan Siaas ### Chat about AI Vicki Reyzelman ### Financial Management Steph Brown ### Humanity Working Paul Slater ### Innovate for Success Daniele Di Veroli ### Jetsoft Pro Coffee Chat Olga Frayt ### Let's Talk Startups Nargis Jafferali and Demos Demetriou ### Managing Innovation John Bessant ### Match Relevant Jake Villareal ### Online Business Startup Tips Thomas Clairmont ### Remotely Possible Adam Riggs ### ScaleUp Radio Kevin Brent ### Secret Ops Podcast Ariana Cofone ### Startup Huddle Jozef Maruscak ### The CTO Show Mehmet Gonullu ### The First Customer Jay Aigner ### The SaaS Podcast Omer Khan ### The Shift Spotlight Ken Paskins & Winter Baserva ### The Slight Edge Advantage Jeremy Torisk ### The Tech Founder's Podcast Marcus Papin ### Think Future Chris Kalaboukis ## Frequently asked questions ### What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. ### What audiences does Zach Rattner speak to? Zach Rattner speaks to executive teams, industry associations, and companies that want AI to produce results rather than pilots. His audiences come mostly from insurance, home services, logistics, financial services, and real estate, where the work is physical and the AI has to survive contact with it. ### Can Zach Rattner tailor a keynote to a specific industry? Yes. Zach Rattner builds the examples in each keynote around the industry and the specific problem the audience is stuck on, rather than delivering a generic version. Name the industry and the problem when you reach out and he will shape the talk around it. ### How far in advance should you book Zach Rattner? Three to six months ahead gives the best chance at a specific date. Zach Rattner can sometimes take shorter turnarounds depending on where he is already traveling that week, so a close-in date is still worth asking about. ### What is included when you hire Zach Rattner to speak? A keynote from Zach Rattner includes a slide deck prepared for that audience, a takeaway attendees keep such as a roadmap or a cheat sheet, and guidance after the event. Workshops also include the materials and templates a team needs to keep going once the session ends. ### What do Zach Rattner's patents cover? Zach Rattner holds 31 granted US patents spanning computer vision, consumer electronics, and telecommunications. Most apply AI to the physical world: measuring interior spaces, streamlining logistics, and making connected devices behave. One of them is a better hair dryer. ### Where is Zach Rattner based and where does he travel? Zach Rattner is based in San Diego, California, and speaks across the United States, Canada, and the United Kingdom. ### What stage setup does Zach Rattner prefer? Zach Rattner presents from the front of the stage rather than behind a lectern, because the connection with the audience is better without furniture in between. A podium is useful for holding a laptop but is not required. ### Does Zach Rattner bring his own presentation equipment? Yes. Zach Rattner runs the presentation from his own MacBook Air in Keynote, which avoids the font and file compatibility problems that come with handing a deck to a venue. He brings his own HDMI dongle and presentation remote. ### What AV setup does Zach Rattner need? Zach Rattner presents in 4k where a venue supports it, with 1080p as the minimum. His slides include video, so the room needs audio as well. --- ## Page: Hands-on AI workshops for enterprise teams | Zach Rattner **URL:** /workshops **Description:** Four hands-on AI workshops led by a working CTO with 31 patents. Compare formats, audiences and pricing, and pick the one your team needs. # Pick the workshop your team actually needs Four hands-on days, each built for a different room. I run them myself, and I still build this for a living. ## Start with the sentence that sounds familiar Every one of these is something a leader has said to me in the first ten minutes of a call. Whichever one you recognize is the workshop your team needs first. - "Every QA vendor is promising us 100% call coverage. Which one do we buy?" Turn Your Call Center Into an AI Asset - "Our AI agents demo beautifully and have been stuck in pilot for six weeks." Mastering Agentic Workflows - "We bought the Copilot licenses. Nothing about how we ship actually changed." Modernize Your Engineering Org for the AI Era - "We keep losing enterprise deals and nobody can tell us it is the security review." The 60-Minute Security Audit™ ## Photographs from recent workshops Recent sessions in Japan, the United Kingdom, Aruba, and across the United States. ## What your team leaves with Each workshop can be delivered in person or virtually and can be tailored to your industry before the day. Prices exclude travel. Booking usually runs three to six months out. ### Turn Your Call Center Into an AI Asset For sales, CS, and operations leaders Format Half or full day Group size 8 to 25 From $10,000+ #### Your team leaves with A decision your team can defend on what to buy, what to build, and what each option locks you into, plus a pilot they can start on. See the full day ### Mastering Agentic Workflows For service, support, finance, HR and operations leaders Format Full day Group size 8 to 25 From $12,000+ #### Your team leaves with Escalation paths, guardrails and audit trails designed against your real processes, a CFO-ready business case, and a 90-day rollout roadmap. See the full day ### Modernize Your Engineering Org for the AI Era For developers, platform engineers and technical leadership Format Full day Group size 8 to 25 From $12,000+ #### Your team leaves with MCP servers wired into your workflows, custom agent skills for QA and code review, and one standard way yourI organization ships with AI. See the full day ### The 60-Minute Security Audit™ For executives in IT, engineering, and operations Format Half or full day Group size 8 to 25 From $15,000+ #### Your team leaves with A vendor inventory, the litmus test itself, a ranked Sleep at Night checklist run on your own operation, and the gap list between where you are and the contracts you want to bid on. See the full day ## Where I have taught and spoken ## Still not sure which one? Tell me what your teams are stuck on and I will tell you which of the four fits, or that none of them do. I would rather say so than sell you the wrong day. Ask me which one fits ## Looking for a keynote instead? Workshops are a working day with one team. If you need a room of several hundred to leave thinking differently, that is a different format and it has its own page. See AI speaking programs --- ## Page: The 60-Minute Security Audit™ | Zach Rattner **URL:** /workshops/60-minute-security-audit **Description:** Turn compliance from a chore into an edge. A proven framework for evaluating vendors and scaling data security protocols that hold up. # The 60-Minute Security Audit™ The four-step method I used to take an AI platform through ISO 27001, SOC 2 Type II, and GDPR, and the contracts that opened. Taught to your executives in IT, engineering, and operations in a language that needs no security background. ## Duration Half-Day or Full-Day ## Format In Person & Virtual ## Group size 8 to 25 people ## Price Range $15,000+ Book Now ## The contracts you cannot bid on are the real cost of weak security. Enterprise buyers, regulated industries, and government programs all gate procurement on the same evidence. If you cannot produce it, the deal never reaches your sales team, and nobody tells you why you were not shortlisted. That is the expensive version of this problem, and it has nothing to do with getting breached. I built Yembo's security posture from nothing and took it through ISO 27001, SOC 2 Type II, and GDPR while shipping product across 20+ countries. I am not a consultant who advises on this. I am the executive who had to pass the audits, and who watched which contracts opened once we did. The 60-Minute Security Audit™ is the method I ended up with, compressed into something your leadership team can run without me and without a security background. Four steps, in order: - Inventory. Every vendor holding your customer data, listed. Most teams find names nobody remembers approving. - The litmus test. A non-technical questionnaire that separates vendors who take security seriously from vendors with a trust page. - Sleep at Night. A prioritized checklist run against your own operation, so the gaps come out ranked rather than as a list of everything. - The posture that sells. What to close, in what order, to reach the compliance bar your target contracts actually require. It is called the 60-Minute Security Audit because step two takes about an hour on any single vendor. The workshop is where your team learns to run all four on their own. ## What your team walks out with - The litmus test itself, a non-technical vendor questionnaire your team runs before handing over customer data. - A completed Sleep at Night checklist, prioritized against the gaps we find in your own operation. - A gap list between where you stand today and the compliance posture that wins corporate and military contracts. - The exact questions to put to your technical teams, written down, so the next review does not need me. ## How the day runs Half or full day, in this order. I tailor the examples to your industry before we start, so the shape holds and the content is yours. - 01 ### Inventory Every vendor holding your customer data, listed in the room. Most teams find names nobody remembers approving. - 02 ### The litmus test The non-technical questionnaire, run live against two or three of your real vendors so the team sees what a good answer and a bad answer look like. - 03 ### Sleep at Night The prioritized checklist against your own operation, producing a ranked list of gaps rather than a list of everything. - 04 ### The posture that sells What to close, in what order, to reach the compliance bar the contracts you want to bid on actually require. - 05 ### Questions for your technical team The exact things to ask, written down, so the next review does not need me in the room. Full-day sessions add a working session on your highest-risk gap. ## Where I have taught and spoken ## I have been through these audits myself Not as an adviser. As the CTO who had to produce the evidence, answer the auditor, and keep shipping product across 20+ countries while it was happening. ISO 27001 Information security management, certified SOC 2 Type II, the report enterprise buyers ask for GDPR Live across European operations ## Run the audit on yourself first The Sleep at Night checklist is the framework this workshop is built on, and it is free. Run it across your own data flows and vendor list before we speak. If it turns up nothing, you do not need me. If it turns up a list, you will know exactly what the day is for. Get the free checklist ## Book this workshop --- ## Page: Mastering Agentic Workflows: AI training for business teams **URL:** /workshops/mastering-agentic-workflows **Description:** A hands-on workshop that turns AI agents into reliable systems for customer service, call centers, finance, HR, and operations teams. # Mastering Agentic Workflows A full day that turns AI agents into dependable systems for customer service, support, finance, HR, and operations ## Duration Full-Day ## Format In Person & Virtual ## Group size 8 to 25 people ## Price Range $12,000+ Book Now ## AI agents that do real work, not demos Every operations leader has seen the demo: an AI agent answers tickets, routes requests, and triages workflows. Then the pilot starts, and six weeks later it is still stuck in proof-of-concept purgatory. Meanwhile the category has moved on without you. Madrona's 2026 Intelligent Applications 40 found an agent owning a specific job in nearly every category it tracks, from IT service management to customer service to marketing operations, and reports that 77 percent of enterprises now re-evaluate their AI vendors at least every six months. Your pilot is on a clock, and what it gets graded on is finished work. The bottleneck is not model intelligence; it is workflow architecture. Moving from a fragile demo to a dependable business system requires clear delegation: knowing exactly what agents handle autonomously, when human-in-the-loop escalation triggers, and how strict compliance guardrails protect customer data. Built for customer service, support, finance, HR, and operations leaders, this intensive workshop turns stalled AI experiments into production-grade systems. We map your actual business processes, implement audit trails, and calculate hard payback metrics so your executive team sees clear financial returns rather than vendor hype. Drawing from 31 patents and real-world enterprise deployments, I give your team the frameworks, vendor scorecards, and a 90-day rollout roadmap to deploy autonomous agents with confidence. Your team leaves with an execution blueprint they put into production the following Monday. ## What your team walks out with - A ranked map of your own operations, support, and finance processes, scored for where agents pay back first. - Escalation paths drawn against your real workflows, showing exactly where an agent hands off to a person. - A guardrail and audit-trail specification your engineering team can build from. - A CFO-ready business case with your deflection rates, handle times, and payback period filled in. - A 90-day rollout roadmap for your top workflow, with owners and checkpoints against dates. ## How the day runs Full day, in this order. I tailor the examples to your industry before we start, so the shape holds and the content is yours. - 01 ### Why the pilot stalled We start with the agent work already underway and find where it stopped. It is almost never the model. - 02 ### Process mapping Your real operations, support and finance workflows on the wall, scored for where an agent pays back first. - 03 ### Drawing the handoffs Where an agent acts alone, where it escalates, and who catches it. This is the part that decides whether it survives contact with customers. - 04 ### Guardrails and audit trails The controls your compliance team will ask about, specified before anything goes live rather than after. - 05 ### The business case Deflection rates, handle times and payback, filled in with your numbers so the CFO conversation is already had. - 06 ### The 90 days Your top workflow, from where it is now to production, with owners against dates. ## Where I have taught and spoken "Zach's insights were practical, engaging, and very well received by the attendees." FIDI 39 Club Osaka, Japan. A working session that closed with a 30-day rollout plan each attendee could start on Monday. ## See where your agents would pay back first If service or support is one of the functions you are automating, start there. My free Call Center Analyzer reads your own conversations and shows what full-coverage analysis surfaces, which is the same evidence we use on the day to rank which processes to automate first. Analyze your conversations free ## Book this workshop --- ## Page: Modernize Your Engineering Org for the AI Era | Zach Rattner **URL:** /workshops/modernize-your-engineering-org-for-the-ai-era **Description:** A hands-on team workshop that turns AI code generation into a standardized engineering system: MCP integration, custom agent skills, and automated QA. # Modernize Your Engineering Org for the AI Era A full day that turns AI code generation into a repeatable operational system your entire engineering organization runs the same way ## Duration Full-Day ## Format In Person & Virtual ## Group size 8 to 25 people ## Price Range $12,000+ Book Now ## Stop playing with AI. Start standardizing how your team ships. AI coding tools are routinely marketed as a magic bullet. Buy the Copilot or Cursor licenses, hand them to your developers, and expect immediate 10x gains. Without standardized systems, what you end up getting is slightly faster typing and a codebase cluttered with sprawling boilerplate. The market has arrived at the same conclusion from the investment side. Madrona's 2026 Intelligent Applications 40 argues that as frontier models commoditize, the value moves to the harness built around them, meaning the layer that decides which model runs, on what context, against which tools, and with what verification. Your competitors can buy the same models you can. The harness is the part nobody can sell you. Or perhaps you solved the infrastructure challenge first. You bought the box and watched Qwen 3.8 generate tokens on your own hardware just to prove it could be done. But running a model is an infrastructure task. Getting eight engineers to ship code differently on Monday morning is a workflow problem. This hands-on technical workshop turns AI code generation into a repeatable operational system run uniformly across your entire organization. We connect MCP servers directly to your internal workflows. We build bespoke agent skills for automated QA, continuous integration checks, and automated code review. And we implement agentic workflows that write, test, and verify code autonomously across frontier cloud APIs, local open-weight models, or hybrid environments. I am the CTO and co-founder of Yembo, the author of Grow Up Fast, and the holder of 31 granted US patents. I have trained over 5,000 technical professionals and engineered computer vision platforms that process millions of production requests globally. This workshop is battle-tested production mechanics, not high-level academic theory. Scope note: this workshop is engineered for software developers, platform engineers, and technical leadership. If your teams want to automate business operations, customer support, HR, or finance instead, that is Mastering Agentic Workflows, the companion workshop built for them. ## What your team walks out with - Agentic design patterns running inside your current sprint cycle, not a slide deck describing them. - MCP servers wired to your internal knowledge bases and operational systems, built during the day. - Custom agent skills for your most expensive QA, code review, and deployment bottlenecks, written and committed. - A prototyping workflow that turns a one-sprint spike into an afternoon, demonstrated on one of your own. - A sandbox configuration for running agents against proprietary information safely. - The same stack pointed at a self-hosted open-weight model, so rate limits and per-token billing stop capping what you automate. - A 90-day roadmap for standardizing the toolchain across every team, with owners against dates. ## How the day runs Full day, in this order. I tailor the examples to your industry before we start, so the shape holds and the content is yours. - 01 ### What your team actually does with the licenses An honest look at how AI is used across your org today, which is usually eight different ways. - 02 ### Agentic patterns in your sprint Plan, execute, verify, applied to a real ticket from your backlog rather than a toy example. - 03 ### Wiring up MCP We connect MCP servers to your internal knowledge bases and operational systems during the session. Hallucinations drop when the model can read your systems. - 04 ### Building agent skills Custom skills for your most expensive QA, code review and deployment bottlenecks. Written and committed by your engineers, not by me. - 05 ### Sandboxes and safety How to give an agent proprietary access without giving it a blast radius. - 06 ### The self-hosted option The same stack pointed at an open-weight model on your own hardware, so rate limits and per-token billing stop capping what you automate. - 07 ### Standardizing A 90-day roadmap for making this one way of working across every team, with owners against dates. ## Where I have taught and spoken “Zach did a fantastic job with our Advanced AI for Developers session. He brought real depth around agentic design patterns and the practical realities of building AI in production – architecture, guardrails, and how to run pilots that actually stick. He tailored the content well to our audience, kept it highly interactive, and our team walked away with concrete frameworks they can apply immediately.” - Pranav Singh, Sr. Director, Software Engineering, Cydcor ## Score your codebase first Run my free readiness checklist over your own repository before we talk. It scores how prepared your codebase is for agentic workflows and custom MCP integrations, and it tells you which of the gaps are worth a workshop day and which your team can close on their own. Score your codebase free ## Book this workshop --- ## Page: Turn Your Call Center Into an AI Asset | Zach Rattner **URL:** /workshops/turn-your-call-center-into-an-ai-asset **Description:** Stop treating your call center as a black box. Unlock the hidden value in your voice data and automate quality assurance with practical AI frameworks. # Turn Your Call Center Into an AI Asset Every QA vendor promises 100% call coverage. The expensive decision is which one to buy, what to build instead, and what each choice locks you into. ## Duration Half-Day or Full-Day ## Format In Person & Virtual ## Group size 8 to 25 people ## Price Range $10,000+ Book Now ## The hard part is not analyzing the calls. It is choosing what to buy. You already know your team reviews about 2% of calls by hand. You know the compliance risks, the coaching moments, and the churn signals in the other 98% are going unread. That part is not in dispute, and every vendor in this category will tell you the same thing before quoting you a subscription. The decision that actually costs money is the next one. Which platform, at what depth of lock-in, on whose model, with your customer audio going where. What to build instead, and what is genuinely not worth building. Which parts of your existing stack already do this and are switched off. Get that wrong and you are three years into a contract that priced itself against your call volume. I have no stake in the answer. I am a CTO who shipped conversation analysis into production, and I hold patents in applying AI to the physical world. I do not resell any of these platforms and I am not going to implement one for you. That is exactly why I am useful in the room for a day. This is a teardown, not a lecture. I get your sales, CS, and operations leaders around one table and we work through your actual call flows, your actual vendor shortlist, and your actual constraints. Your team leaves able to: - Judge a QA platform on the three things that decide the contract, rather than on the demo. - Tell which parts of full-coverage analysis to buy and which to build. - Extract revenue-saving intelligence from raw conversations without a machine learning team. Full coverage is table stakes now. The advantage is in choosing well, deploying it without stalling, and freeing your people for the work that needs a human. That only happens when leadership agrees on the decision before the procurement cycle starts. ## What your team walks out with - A scored shortlist of the QA platforms you are actually considering, with the build-versus-buy call made - A written detection map naming which churn signals, compliance risks, and coaching gaps to flag first in your own call data - A vendor evaluation checklist your team can run on the next platform without me in the room - A pilot plan with a named owner, a start date, and the number it has to hit to be judged a success - An operational playbook for moving your QA team off manual auditing and onto retention work ## How the day runs Half or full day, in this order. I tailor the examples to your industry before we start, so the shape holds and the content is yours. - 01 ### Where the 98% is going We open on your own numbers. Volume, current sampling rate, what QA costs you today, and which failures are getting through. - 02 ### The vendor teardown Your actual shortlist on the whiteboard, scored against the terms that decide a contract rather than the demo. - 03 ### Build, buy, or already own it What is worth building, what never is, and which capability is sitting switched off in a tool you already pay for. - 04 ### Your detection map Which churn signals, compliance risks and coaching gaps to flag first, ranked against what they cost you. - 05 ### The pilot A named owner, a start date, and the single number that decides whether it worked. Full-day sessions run this against real call data. ## Where I have taught and spoken "Zach's insights were practical, engaging, and very well received by the attendees." FIDI 39 Club Osaka, Japan ## Start with your own calls Before you book anything, run a sample of your own conversations through my free Call Center Analyzer. It takes minutes, it costs nothing, and it shows you what full coverage surfaces that a 2% sample misses. Most people arrive at the workshop already knowing what they want to fix. Analyze your calls free ## Book this workshop --- ## Page: Grow Up Fast: Lessons from an AI Startup | Zach Rattner **URL:** /shop/grow-up-fast **Description:** Get the inside story of scaling an AI startup at Yembo. This must-read book provides strategic insights for tackling uncertainty and cultivating team agility. ## Ratings & Reviews ### Jason Z (5.0) #### A Must-Read Management Gem: Humor, Wisdom, and Practical Insights Such a great read! This book is funny, lighthearted, and brimming with valuable advice. If you're in any management role, this is a must-read. It's deeply encouraging and leaves you with a heightened sense of determination to enact positive changes. The real-life stories illustrate simple, actionable steps that anyone can take to benefit their employees or company. I've already started embracing many of the suggestions. The author's approach to breaking down problems and addressing challenges is both remarkable and rooted in simple, common sense. It really makes you think about how you are currently doing things. The chapters are concise and filled with fun illustrations and graphics, ensuring readers remain engaged. I don't think I've ever read through a business book so quickly. It's earned a spot on my bookshelf right next to classics like "Good to Great" and "The 7 Habits of Highly Effective People." March 17, 2026 (USA) ### Steven Quintana (5.0) #### Amazing and thought provoking life lessons from the start-up founder perspective! I just finished reading Grow Up Fast and cannot recommend it enough! The book peels back the curtain into Zach's journey from his time as a Software Engineer at Qualcomm, to the beginnings of his company Yembo and where it is today. Zach writes Grow Up Fast in a way that is easy to digest while presenting challenging and thought provoking ideas, with a bit (maybe more than a bit) of his humor sprinkled in. Personally, as someone who has switched careers to software development, I loved getting to hear stories from employees about their journeys. It's clear from their testimonials that an awesome culture has been built and something right is going on over at Yembo! If you are working in a start-up looking for lessons learned through trial and error, a manager looking for fresh perspective, a new or seasoned Software Engineer, or just anyone looking for some practical life lessons (hint: 'The Two Minute Rule'), then this is definitely a book to read in the new year! March 17, 2026 (USA) ### Phillip Hunter (5.0) #### Insightful and thoughtful A worthwhile exploration of the journey from deciding to pursue the dream of creating a company through the details large and small to make it reality. Fast-paced, humble, and contemporary. A good mixture of founder story, lessons learned, personal growth, and pragmatic idealism. March 17, 2026 (USA) ### William Fr (5.0) #### Leadership, Growth and Resilience "Grow Up Fast" is a game-changer guide for both aspiring and seasoned leaders striving to grow and navigate life's business challenges with resilience and determination. The book is filled with ideas and insights on leadership, personal growth, teamwork, and overcoming adversity. As I read, I couldn't help but take notes — over 80 key ideas that I'm excited to revisit and apply to my own business journey... March 17, 2026 (USA) ### Nadia Groen (5.0) #### Worth reading This book is worth your time. I read it a few months ago and still go back to it. As a newly founder myself, I learned some practical ways to look at company growth, team, and mindset. March 17, 2026 (USA) Question. Learn. Adapt. # Grow Up Fast: Lessons from an AI Startup (Hardcover) $29.99 USD Add to Cart The untold narrative of our era is that there are yet untraveled paths to tread and new discoveries to uncover in the world of artificial intelligence (AI). In Grow Up Fast: Lessons from an AI Startup, Zach Rattner, an entrepreneur who journeyed from being a corporate employee to a startup founder, reveals how we can navigate this relatively unknown terrain to create novel leadership and management solutions. Zach, CTO and Co-Founder of the AI startup Yembo, offers an honest and enlightening perspective on the journey of building a startup in the rapidly evolving field of AI. This book isn't about quick success or easy wins; instead, it emphasizes the importance of adaptability, patience, and resilience in the face of unexpected hurdles. It's a guide for those who are eager to venture into the world of AI, based on Zach's own trials and triumphs. Zach opens with the idea that in an AI startup, uncertainty and discomfort are not hurdles but the thing that makes you grow. While many are dazzled by the newest technology and the pace of it, Zach argues that progress should not be confined to established tech hubs or pre- defined paths. Every industry, every business holds the potential for evolution. It is all rooted in one crucial capability that every leader must cultivate: the power to question, learn, and adapt. The book includes key lessons learned: - The importance of questioning assumptions - The value of diversity within a team - The traps and benefits of feedback - The beauty of constraints Zach shows us that getting comfortable with discomfort, managing uncertainty well, and trusting teams are critical elements in a startup's progression. Grow Up Fast brings forth both an invigorating vision of future growth in the AI sector, and a fresh approach to understanding innovation: it all begins by asking the right questions which then lead you to discover untapped potential. Zach's experiences and insights serve as an inspirational compass for those eager to embark on their own entrepreneurial journey in the captivating yet intricate world of AI startups. ## Product details - ASIN : B0C9SF8JXJ - Publication date : July 6, 2023 - Language : English - Print length : 276 pages - Item Weight : 11.4 ounces - Dimensions : 5.5 x 0.63 x 8.5 inches ### Grow Up Fast: Lessons from an AI Startup (Hardcover) The untold narrative of our era: how to build an AI startup properly and efficiently. $29.99 USD Add to Cart ## Resources for your AI journey ### Book Zach to speak Turn AI into ROI with practical strategies. Zach has trained over 5,000 professionals worldwide. LEARN MORE --- ## Page: AI Conversation Starter Kit | Zach Rattner **URL:** /shop/ai-conversation-starter **Description:** Turn AI buzz into prioritized pilot projects. Access a tactical framework designed to translate operational pain points into actionable AI initiatives. ## Ratings & Reviews ### Phillip Hunter (5.0) #### Insightful and thoughtful A worthwhile exploration of the journey from deciding to pursue the dream of creating a company through the details large and small to make it reality. Fast-paced, humble, and contemporary. A good mixture of founder story, lessons learned, personal growth, and pragmatic idealism. March 17, 2026 (USA) Turn AI buzz into prioritized pilot projects. # AI Conversation Starter $29.99 USD Add to Cart ## Description Turn AI buzz into prioritized pilot projects. Many organizations recognize the need for an AI strategy but struggle to move past abstract brainstorming. The AI Conversation Starter is a practical workshop framework designed to bridge the gap between high-level discussions and actionable, business-driven initiatives. Developed by AI founder and author Zach Rattner, this framework provides a structured series of questions that guide teams to identify real operational pain points and translate them into practical AI applications. Rather than starting with the technology, this tool forces teams to define success metrics, identify operational constraints, and address potential risks upfront. Key Features: - Structured Discovery: A guided question set divided into strategic phases, starting with core organizational challenges and ending with pilot execution. - Risk & ROI Alignment: Prompts your team to address critical business factors—including ROI measurement, data security, and stakeholder alignment—before committing resources. - Actionable Outcomes: Leaves your group with a refined short list of pilotable AI projects and clear, measurable next steps. - Expert-Designed: Built on real-world lessons from scaling enterprise AI, filtering out industry hype to focus on practical utility. Who is this for? Leadership teams, department heads, and project managers who need to implement AI solutions that deliver clear business value and address concrete operational inefficiencies. ## Resources for your AI journey ### Book Zach to speak Turn AI into ROI with practical strategies. Zach has trained over 5,000 professionals worldwide. LEARN MORE --- ## Page: Grow Up Fast: Lessons from an AI Startup by Zach Rattner **URL:** /grow-up-fast **Description:** Ambitious projects start with uncertainty. Manage the chaos, validate ideas quickly, and build a team you trust, from CTO and inventor Zach Rattner. # GROW UP FAST Lessons from an AI Startup ORDER YOUR COPY NOW ## Ambitious projects start with uncertainty. The most successful leaders learn to manage the chaos and grow up fast. ## Embarking on the entrepreneurial journey is a thrilling, yet daunting endeavor. GROW UP FAST serves as your companion on this journey. ORDER YOUR COPY NOW ## About the Book Zach shows us the lessons he learned along the way as he journeyed from an employee mindset to a founder mindset. The lessons he learned are universal amongst entrepreneurs, but Zach gives us a unique spin, as his startup is in one of the fastest growing areas in the business world – artificial intelligence. Join Zach in GROW UP FAST as he explores the thrilling highs and lows of an AI startup journey. Learn to navigate uncertainties, validate ideas quickly, and build a trusted team. Test your assumptions Feedback is dangerous Constraints are beautiful Harmony not homogeneity Make yourself uncomfortable Manage uncertainty Empower others ## About the Author ### Zach Rattner Zach Rattner, the Chief Technology Officer and co-founder at Yembo, is a software engineer with over 15 years of experience. He has a B.S. in Computer Engineering from Virginia Tech. Yembo is the leader in AI-powered virtual surveys with over 5 million videos inspected and customers in over 20 countries. Zach's previous projects include building a flashcard studying tool that scaled to over a million users and serving as the software lead for Qualcomm's internal innovation program, Qualcomm ImpaQt. He is in the top 2% of contributors on Stack Overflow and has 31 granted US patents with several more pending. GROW UP FAST is his first book. He has three children and lives in California with his wife Lindsay. ORDER YOUR COPY NOW ## Order Now ### Amazon Order Now ### Apple Books Order Now ### Google Play (English) Order Now ### Google Play (Spanish) Order Now ### Spotify Order Now ## Who this book is for - Experienced professionals new to AI - Managers leading multidisciplinary teams - New or prospective startup founders - Early career tech professionals - Students Here for the audiobook supplement? Download Now the Grow Up Fast audiobook supplement PDF --- ## Page: 60-Minute Security Audit™ checklist | Zach Rattner **URL:** /resources/60-minute-security-audit **Description:** Protect your data when adopting AI. Download the 60-Minute Security Audit blueprint and test vendor compliance without the jargon. # 60-Minute Security Audit™ ## A leader's protocol for vetting AI vendors Your software vendors are either an asset or a liability. There is no middle. Integrating AI isn't just about efficiency; it's about access. When you deploy third-party AI, you grant outside code access to your proprietary data. If that code is weak, your reputation is the collateral. I am Zach Rattner, CTO and Co-Founder of Yembo. Our AI platform is deployed in 20+ countries, enabling businesses to scale to 3x their daily volume. I hold 31 granted US patents. I do not deal in theory. I build secure, scalable technology for the most regulated industries on earth. ### Engineering trust, not marketing claims Global insurance carriers and government agencies do not accept "we're secure" as an answer. They demand proof. We engineered Yembo's codebase from day one to exceed the world's strictest compliance standards: - ISO 27001 - SOC 2 - GDPR - NIST 800-171 I know the grit required to build a fortress around sensitive data while running a high-growth company. Protecting your business does not require a computer science degree. It requires a founder's perspective on risk. ### The non-technical litmus test Not knowing exactly where your data goes once it hits a vendor's server creates "Black Box" anxiety. This uncertainty constantly stalls enterprise AI adoption. I created the 60-Minute Security Audit™ as a practical framework for executives. The premise is simple: if a software vendor has strong security practices built-in, their team should be able to complete this audit in less than one hour. It filters out high-risk vendors immediately, ensuring your AI strategy is a driver of ROI, not a liability. Question 01 ### Do you hold a current SOC 2 Type II or ISO 27001:2022 certification? If so, is there a dashboard we can use to follow real-time compliance throughout the year? These certifications prove an independent auditor has verified the vendor's security controls over time, not just in a one-off check. The companies that take this seriously run a trust center you can check yourself, showing live attestation against their controls. 🚩 Red Flag: "We follow SOC 2 principles" (but have no report) or relying solely on their cloud provider's (e.g., AWS) security. The ISO 27001 standard was refreshed in 2022. In 2026, there are risks if a company is still adhering to the older 2013 standard. Question 02 ### Do you carry dedicated cyber liability insurance? General business liability insurance often excludes cyber incidents. If a vendor causes a breach, you need to know they have a specific policy to cover forensic investigations and lawsuits rather than going bankrupt and leaving you with the bill. 🚩 Red Flag: "Our general liability policy covers it" or coverage limits under $1M. Question 03 ### Is customer PII (Personally Identifiable Information) and video data encrypted at rest and in transit? Without this, your customer's names, addresses, and video inventories are readable by anyone who gains access to the hard drive or intercepts the Wi-Fi connection. 🚩 Red Flag: They mention "encryption" generally but can't name the specific algorithms (e.g., AES-256, TLS 1.3). Question 04 ### Is your platform protected by a firewall? Without a properly configured firewall, your platform is essentially sitting on the open street with the front door wide open to automated "bot" attacks and hackers. 🚩 Red Flag: No regular testing schedule for firewall configurations. Question 05 ### Do you enforce Multi-Factor Authentication (MFA)? Passwords are easily stolen or bought on the dark web. If a vendor's support staff can access your data without MFA, your data is one "phishing" email away from being exposed. 🚩 Red Flag: "We encourage it but don't enforce it" or "Only for admins." Question 06 ### Can you provide your most recent penetration test summary report? Penetration testing is like a fire drill for your digital security. It involves hiring ethical hackers to intentionally try to break into the system to find weaknesses before real criminals do. A report from 2023 or older is useless in 2026. 🚩 Red Flag: The vendor has no report, refuses to share a summary report, or the report is more than 12 months old. Another red flag is if the report shows "Critical" or "High" vulnerabilities that remain unresolved months after the test. Question 07 ### Do you maintain a formal process such as a Vulnerability Disclosure or Bug Bounty Program for external vulnerability reporting? Even the best internal security teams can miss vulnerabilities. Ethical hackers often discover these blind spots in the wild. Without a dedicated reporting channel, critical warnings get lost in general support queues or well-meaning researchers face legal threats. A Vulnerability Disclosure or Bug Bounty Program ensures these external findings are safely reported and patched before attackers can exploit them. 🚩 Red Flag: Directing reports to a general support email, claiming internal testing catches everything, or lacking a safe harbor policy to protect well-intentioned researchers. ### Keep reading Enter your email to unlock the full article. By submitting your email, you consent to be contacted in accordance with our Privacy Policy. ## Bring the AI security blueprint to your team Book Zach for a high-impact workshop. Get actionable, ROI-driven frameworks, not theory. Trusted by enterprise leaders in insurance, logistics, and tech. Check Availability Close Modal ### Secure your organization with the full blueprint. Get the full 30-question blueprint, including the specific AI & data framework I use at Yembo. Use it to audit your vendors in under 60 minutes. By submitting your email, you consent to be contacted in accordance with our Privacy Policy. --- ## Page: Agent-Ready Codebase Audit | Zach Rattner **URL:** /resources/agent-ready-codebase-audit **Description:** Download the Agent-Ready Codebase Audit. A deterministic 10-point checklist for your platform's readiness for MCP integrations and AI agents. # Agent-Ready Codebase Audit Too often, software teams treat AI coding tools as a magic bullet. They buy the licenses, hand them to their developers, and expect immediate 10x gains. But without strictly standardized systems in place, the result is just slightly faster typing, fragmented workflows, and a codebase cluttered with hallucinated boilerplate. Building reliable, autonomous AI requires more than just academic theory — it requires battle-tested engineering. As CTO and Co-Founder of Yembo, I know what it takes to scale AI reliably because I've built an AI startup from zero to global deployment. My team and I have built AI-powered computer vision platforms used daily in over 20 countries, turning complex artificial intelligence into tangible, physical-world results for enterprise industries. Along the way, I've racked up 31 granted US patents and trained over 5,000 professionals worldwide on how to safely deploy AI into production. I don't teach high-level academics; I teach production-ready best practices. The Agent-Ready Codebase Audit is born directly from this hands-on experience. It strips away the hype and provides a deterministic, 10-point framework to help you evaluate your codebase's readiness for custom MCP integrations and autonomous agents. Enter your email below to get the free guide and find out exactly what foundational gaps your engineering team needs to close before you scale. ## A 10-point framework to stop playing with AI and start using it 1. Do you follow standardized, predictable processes from ticket creation to implementation, testing, and release? - Why it's important: AI agents thrive on predictability. If your human developers don't have a standard way of working, agents won't either. Without standardized systems, introducing AI will result in dead-ending workflows and wasted tokens. - How to get ready: Audit your Agile or Kanban workflows. Create strict, mandatory templates for bug reports and feature requests in tools like Jira or Linear so that every task follows a predictable lifecycle. 2. Are your ticket requirements and "Definition of Done" defined and documented? - Why it's important: Agents cannot read minds or make intuitive leaps about business logic. If requirements are vague, the agent will fill the gaps with guesses, leading to a codebase cluttered with hallucinated boilerplate. - How to get ready: Train product managers and tech leads to write hyper-specific acceptance criteria. If a junior developer couldn't build it based only on the ticket text, an agent definitely can't. 3. Do you have separated environments for development, staging, and production? - Why it's important: Agents will make mistakes. You need strict, sandboxed guardrails to safely transition your team into using them. They need a safe playground to break things without taking down live customer data. - How to get ready: Stop testing in production. Set up distinct, isolated environments (e.g., dedicated Docker containers or cloud instances) where agents can safely deploy and test code. 4. Do you have comprehensive automated tests (unit, integration, end-to-end)? - Why it's important: You cannot manually review every line of code an agent writes at scale. Automated tests are the primary defense mechanism to catch and eliminate dangerous code hallucinations before they get merged. - How to get ready: Pause feature development if necessary and pay down testing debt. Establish a baseline of test coverage for your critical paths and enforce rules that no code gets merged without passing tests. 5. Are your deployments and releases fully automated (CI/CD)? - Why it's important: To build a standardized, AI-native engineering machine , agents must be able to autonomously plan, execute, and verify complex architecture. If a human has to manually click "deploy" or move files over FTP, you bottleneck the agent's speed. - How to get ready: Implement CI/CD pipelines (like GitHub Actions, GitLab CI, or CircleCI) that automatically build, test, and deploy code when changes are pushed. 6. Are your internal APIs clearly structured and documented? - Why it's important: To give agents "skills," you need to connect AI directly to your APIs. This is often done by leveraging Model Context Protocol (MCP) to set up custom agent skills. Unstructured APIs mean agents can't interact with your systems. - How to get ready: Adopt standardized API documentation, such as OpenAPI/Swagger specifications, for all internal and external endpoints. 7. Have you clearly identified your code-review and QA bottlenecks? - Why it's important: The highest ROI for agents isn't just writing code; it's automating your most expensive QA, code-review, and deployment bottlenecks. You need to know where these bottlenecks are to deploy agents effectively. - How to get ready: Measure your team's cycle times. Look at how long PRs sit waiting for review or how much time is spent on manual QA, and target those areas for your first agentic pilots. 8. Is your system architecture reasonably modular or decoupled? - Why it's important: Agents struggle to navigate massive, tightly coupled spaghetti code monoliths because the context window required to understand the ripple effects is too large. - How to get ready: Begin refactoring large monoliths into smaller, distinct modules, services, or bounded contexts with clear separation of concerns. 9. Do you have robust error tracking and system observability in place? - Why it's important: When an agent pushes code that breaks something in production, you need to know exactly what broke and why, instantly. You cannot rely on users to report bugs created by AI. - How to get ready: Implement tools like LogRocket, Sentry, or Datadog to capture real-time errors, performance metrics, and user session data. 10. Is your team culturally ready and trained to collaborate with AI? - Why it's important: Tools don't transform organizations; people do. If your team views AI as a threat or a fad, adoption will fail. They need to understand how to prompt effectively, review AI code, and trust the new workflows. - How to get ready: Invest in comprehensive training. Build internal playbooks on AI best practices, and celebrate early wins to foster a culture of curiosity and continuous improvement. ## Ready to take things to the next level? I run full and half-day workshops on readying your codebase for AI agents. Learn More about the Modernize Your Engineering Org for the AI Era workshop --- ## Page: Call Center Analyzer AI tool | Zach Rattner **URL:** /resources/call-center-analyzer **Description:** Stop letting customer insights go to waste. Implement the Call Center Analyzer to scale your QA with AI and identify exactly why jobs aren't booking. # Call Center Analyzer ## Stop guessing why jobs aren't booking: scale your call center QA with AI Every call your team takes is a goldmine of data, holding the exact reasons a customer books a job or hangs up to call a competitor. But if your sales managers are relying on manual processes to extract those insights, you are moving too slowly. Right now, quality assurance teams are drowning in transcripts, struggling to give your sales reps the timely, targeted coaching they need to overcome objections and close more deals. That manual bottleneck actively limits your booking rate and adversely affects your revenue. That's why I built the Call Center Analyzer. I'm Zach Rattner, CTO and Co-Founder of Yembo. My team and I build AI products that are used daily by moving and insurance companies across 20+ countries. We've helped organizations complete 3x more inspections per day while actively improving their customer experience. I currently hold 31 granted US patents, and if my experience bringing AI to the physical world has taught me anything, it's this: AI only matters when it drives tangible ROI in the real world. This secure, AI-powered application is explicitly designed to automatically grade call center transcripts. By eliminating the tedious task of manually reading and scoring every interaction, this tool frees up your sales leads to do what they do best: mentor their teams, refine their pitches, and drive meaningful revenue. ### What makes this tool different? - Instant Sales QA: Ensure accurate, consistent feedback on every call, instantly, so you can scale your QA process without adding operational overhead. - Built for Moving: Whether your reps are quoting interstate moves, local, international, or national accounts, this tool helps you identify exactly where they are dropping the ball on the script. - Actionable Coaching, Zero Fluff: Get straight to the insights that matter. Quickly identify your top performers and pinpoint exactly who needs help handling pricing objections. - Secure From Day One: Technology only matters when it makes lives better, and it has to be secure from the start. We built this with essential safeguards out of the box, ensuring your sensitive customer data stays completely locked down. It's time to stop playing with AI and start building with it. The Call Center Analyzer gives you a clear, deployable blueprint so your sales teams can work smarter, close faster, and give every customer a better experience. Try out the free tool at https://callcenteranalyzer.com ## Ready to take things to the next level? I teach full and half-day courses on leveling up your call center with AI tools. Learn More about the Turn Your Call Center into an AI Asset workshop --- ## Page: Apple silicon AI cluster: M4 to M6 local LLM blueprint **URL:** /projects/ai-mac-cluster **Description:** How we cut $40k a year of cloud AI spend with an Apple silicon cluster, and how the M5 Ultra, M5 Pro, and M6 change the build today. # How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year ## Real-world infrastructure blueprints from a CTO deployed in 20+ countries. A local server cluster utilizing M4 and M4 Pro Mac minis to run AI workloads locally, reducing the need for costly cloud services. See AI Speaking Programs ## Executive Summary - The Challenge: Runaway cloud costs for high-volume AI speech transcription. - The Solution: A localized, on-premise Apple silicon architecture. - The Result: Reduced Google Cloud spend by $40,000 annually while maintaining ISO 27001 / SOC 2 compliance. As the CTO of Yembo, where our AI platform processes data across 20+ countries, I am constantly auditing our tech stack for efficiency. I moved a major enterprise workload over to a local M4 Mac mini cluster to prove that scaling AI doesn't have to mean scaling your cloud bill. This move eliminated our reliance on Google Speech to Text. At the time, that service cost $0.016 per minute. The Macs are using whisper.cpp, which runs on the Neural Engine and GPUs in the Apple silicon to transcribe calls locally. Transcription requests come in via SQS, and there's an autoscaler on Kubernetes in AWS that idles at zero, ready to pick up the work if there were to be an outage. The performance is incredible: a single M4 Pro can keep up with 20 concurrent calls at 2x realtime. It's truly a testament to what these machines can do. However, speech transcription is just the beginning—only two of the eight machines in the cluster are dedicated to AI. ### Speech Transcription The original two AI services that started it all. Using whisper.cpp and Silero VAD, these dedicated nodes replaced Google Speech to Text. ### GitHub Action runners Paired with Biome, repatriating our CI/CD pipeline dropped our full-repo build and lint times from four minutes down to just 40 seconds. ### CircleCI Self-hosted runners specifically configured to accelerate native app builds, capitalizing on the performance leaps of Apple silicon vs x86. ### Playwright Automated QA Our heavy daily regression testing suite is executed via self-hosted GitHub Action runners, keeping the tests fast and avoiding expensive cloud execution time. ### Architecture & Specs - M4 Pro Mac minis handling local AI inference - whisper.cpp + Silero VAD for transcription - SQS for request queuing - AWS Kubernetes autoscaler (idling at zero) for fallback - Handles 20 concurrent calls at 2x realtime per machine ### Enterprise Compliance My company is ISO 27001:2022 and SOC 2 compliant, so getting the details right to be able to launch this was a bit of a project. The cluster adheres to strict security and compliance requirements while keeping inference localized. ## The Apple silicon advantage: unified memory for local LLMs Running large language models in the cloud usually requires renting expensive enterprise-grade GPUs with dedicated VRAM. The M4 and M4 Pro Mac minis change that math with their Unified Memory Architecture. By sharing one large pool of high-bandwidth memory between the CPU and the GPU, a single M4 Mac mini can load and run models that would otherwise fail on consumer hardware. Hardware Configuration Model & Quantization Framework Performance M4 Pro with 64 GB Unified Memory Llama 3 8B quantized to Q8_0 llama.cpp and Ollama 58 tokens per second M4 Pro with 64 GB Unified Memory Llama 3 70B quantized to Q4_K_M llama.cpp and Ollama 14 tokens per second Base M4 with 24GB Unified Memory Llama 3 8B quantized to Q4_K_M MLX Framework 42 tokens per second To size a configuration yourself, the Apple silicon local LLM memory and speed calculator estimates required unified memory, prefill speed, and generation tokens per second for every chip from M1 to M6. This cluster was built on M4 and M4 Pro hardware, and the architecture has outlived the chips in it. If you are specifying a build today, the decision has moved to whether you buy one M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of M6 Mac minis, which is the memory bandwidth argument laid out in M5 Ultra vs. M5 Pro vs. M6 for local AI. The models have moved too: a 27B class coding model like Qwen 3.8 running locally under Ollama now does the work that justified a cloud API contract when this cluster went into service, and the full field of current options is weighed in the best local LLMs for agentic coding. Which runtime serves those weights turns out to matter too: benchmarking MLX against Ollama on identical weights put a 40 percent gap between them on one model and reversed it on another. ## Building a private AI agent appliance With the rapid rise of autonomous agent frameworks like OpenClaw and the Hermes Agent, the need for a highly secure, private runtime environment is critical. Deploying these agents locally on our M4 cluster prevents proprietary enterprise data, internal communications, and database schemas from being transmitted to third-party APIs. Our cluster functions as a highly secure private AI appliance. Since all model inference is executed within our restricted local network perimeter, we eliminate external data transit entirely. This architecture allowed us to easily pass our rigorous ISO 27001:2022 and SOC 2 audits, showing that local AI can be both highly innovative and structurally compliant. ## Why this matters for business leaders AI doesn't have to mean runaway cloud bills. By strategically offloading specific, high-volume workloads like transcription to specialized, cost-effective on-premise hardware like Apple silicon, businesses can achieve massive ROI while maintaining enterprise-grade reliability and security compliance. This is no longer a contrarian position. Sequoia Capital's guide to sovereign AI, Own Your Intelligence, makes the same case from the investor's side of the table: as open-weight models close the gap with frontier APIs, owning your intelligence layer protects margins, keeps proprietary data internal, and turns the model itself into product differentiation. The cluster on this page is what that strategy looks like as hardware on a shelf. The enterprise side is moving the same way. In August 2026, Thomson Reuters launched Thomson, an in-house large language model, alongside an upgraded CoCounsel Legal platform, and Simply Wall St asked the obvious question: is Thomson Reuters building a defensible edge by owning its own legal AI? Their answer is that owning the model, rather than renting one, is what turns proprietary content into a moat competitors cannot rent their way around. Thomson Reuters runs that model with cloud partners rather than on hardware it owns, which is the part this series takes further: the strategic argument for owning the model is the same argument for owning the machine it runs on, and at Apple silicon prices that second step is now within reach of teams far smaller than Thomson Reuters. ### Bring your AI strategy down to earth. If you want a proven, actionable blueprint to manage cloud costs, optimize your hardware, and securely deploy enterprise AI without the hype, let's talk. See AI Speaking Programs ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview Currently Reading ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Current Page Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ## Community Discussions The concept of using Apple silicon for localized AI infrastructure resonated strongly with the developer and self-hosting communities. You can read the original case studies and follow the deep-dive technical discussions here: - Reddit: r/LocalLLaMA - Reddit: r/mac - Reddit: r/selfhosted --- ## Page: M5 Ultra vs. M5 Pro vs. M6: which Mac for local LLMs? **URL:** /projects/ai-mac-cluster/m5-ultra-vs-m5-pro-vs-m6 **Description:** M5 Ultra Mac Studio, M5 Pro Mac mini, or a swarm of M6 Mac minis for local LLMs? The bandwidth math behind tokens per second, and the buying call. # M5 Ultra vs. M5 Pro vs. M6 for local AI Apple just announced its first 2nm chip in an $899 box, a quad-die monster with 512GB of unified memory, and an M5 Pro Mac mini sitting between them. Which one belongs in your local AI stack? The answer depends on one number most spec sheets bury. Read on for the math. ## The new era of Apple silicon AI Apple does not usually give AI engineers two interesting decisions in the same product cycle. This time it introduced the M6 and the M5 Ultra together. The M6 brings Apple's first 2nm process to the Mac mini, the cheapest box in the lineup. The M5 Ultra brings the first quad-die M-series chip to the Mac Studio, with a memory ceiling that used to require a server rack. So every engineer building local AI infrastructure now faces the same fork in the road: build a swarm of cheap M6 Mac minis and distribute the work, buy a single high-bandwidth M5 Ultra Mac Studio and keep everything in one memory pool, or split the difference with an M5 Pro Mac mini that runs one serious model on one desk. I run a rack of Mac minis in production, so I have real affection for the swarm. But affection is not a benchmark. Let's work the problem. ## Spec showdown: the AI-centric hardware matrix Ignore the consumer metrics. Gaming frame rates and video export times tell you nothing about how a chip serves a language model. For local LLM work, three things decide your experience: how much memory the chip can address, how fast it can read that memory, and how quickly it can chew through a long prompt. Here is how the three chips compare on the specs that actually matter. Spec M6 Mac mini M5 Pro Mac mini M5 Ultra Mac Studio Process 2nm, Apple's first 3nm-class 3nm-class, quad-die Package Single die Single SoC Four dies joined by UltraFusion, a first for the M series CPU 12 cores: 2 super, 4 performance, 6 efficiency Up to 18 cores Up to 36 cores: 12 super, 24 performance GPU 12 cores, each with a Neural Accelerator Up to 20 cores, each with a Neural Accelerator Up to 80 cores, each with a Neural Accelerator Neural Engine Dual 16-core, up to 2x peak compute over prior generations 16-core 32-core Peak AI compute uplift Up to 30% over M5 Up to 4x over M4 Pro Up to 4.5x over M3 Ultra Unified memory ceiling 32GB 64GB 512GB Memory bandwidth Up to 170GB/s, about 10% over M5 Up to 307GB/s 1.2TB/s, about 50% over M3 Ultra Starting price $899 $1,699 $5,499 Availability September 22, 2026 September 22, 2026 September 22, 2026, but 512GB not until late October Role in a local AI stack Entry-level swarm node Single-box workhorse for one mid-size model Single-machine workstation for frontier-scale models ### M6: the 2nm swarm node The M6 is not trying to run a 400-billion parameter model, and it shouldn't apologize for that. What it is trying to be, and what it succeeds at being, is the best cheap node ever made for distributed inference. The 2nm process buys efficiency, the Dual 16-core Neural Engine buys concurrent small-model throughput, and 32GB of unified memory at 170GB/s is enough to serve a mid-sized open-weight model per node. If your plan involves Exo or llama.cpp RPC across a rack, this is the box you buy five of. ### M5 Pro: the one-box workhorse The M5 Pro is the middle path, and for a lot of teams it is the honest answer. In the 2026 Mac mini it pairs up to 64GB of unified memory with 307GB/s of bandwidth, which runs a 27B class model with real headroom and squeezes a 70B model at 4-bit quantization with a short context. One developer, one desk, one serious model: that is the M5 Pro's territory, and the LLM speed calculator will show you exactly where its 64GB ceiling starts to pinch. ### M5 Ultra: the quad-die monolith The M5 Ultra is Apple's first quad-die system-on-chip, four dies fused into one address space by UltraFusion. The headline for AI work is the memory system: up to 512GB of unified memory at 1.2TB/s. That is enough capacity to load a model in the hundreds of billions of parameters entirely locally, and enough bandwidth to generate tokens at a pace that feels like a hosted service. This is the machine you buy when the job used to require an enterprise GPU workstation and a procurement committee. ## What the silicon changes mean for AI developers Spec tables are nice, but the interesting question is what these numbers do to your workload. Three architectural changes matter here, and they map cleanly onto the three phases of serving a language model. ### GPU neural accelerators and prompt processing All three chips embed a Neural Accelerator inside each GPU core. That placement is the whole story. The Neural Engine has always been fast, but it lives on the far side of the chip from where LLM inference actually runs. Putting dedicated matrix-multiply hardware inside the GPU cores means the acceleration lands exactly where frameworks like MLX and llama.cpp already do their math. Apple's numbers: up to a 30% peak AI compute increase on the M6 over the M5, and up to 4.5x on the M5 Ultra over the M3 Ultra it replaces. The Ultra's multiplier is bigger because it compounds two effects, more GPU cores and an accelerator inside every one of them. The M5 Pro carries the same per-core accelerators the rest of the M5 generation introduced, scaled to its 20 GPU cores. Where you feel this is prompt processing, the prefill phase. Before a model writes its first token, the GPU has to process every token already in the context window, and that pass is compute-bound. Agentic workflows re-read a large context on every single turn, so prefill speed is the difference between an agent that feels responsive and one you alt-tab away from. The Neural Accelerators attack exactly that bottleneck. ### The memory bandwidth gap: 170GB/s vs. 1.2TB/s Prefill is compute-bound. Generation is not. Once the model starts writing, every new token requires streaming the active weights back out of memory, so tokens per second is capped by memory bandwidth divided by the size of the weights being read. This is the math I wish every spec sheet led with, because it turns a marketing number into a speed you can feel. Take Qwen3.8-27B, the model this series recommends for local agentic coding. The default 4-bit quant is 18GB of weights. Divide each chip's bandwidth by 18GB and you get the theoretical ceiling on generation speed: Chip Bandwidth Theoretical Ceiling Realistic Range M6 170 GB/s 170 ÷ 18 ≈ 9.4 tokens/s Roughly 6 to 8 tokens/s M5 Pro 307 GB/s 307 ÷ 18 ≈ 17 tokens/s Roughly 10 to 14 tokens/s M5 Ultra 1,200 GB/s 1,200 ÷ 18 ≈ 67 tokens/s Roughly 40 to 55 tokens/s Realistic ranges assume the 60 to 80 percent of theoretical bandwidth that inference runtimes typically sustain. Longer contexts pull the numbers down further on both chips. Read that table twice, because it settles the single-machine question. Six tokens per second is a model you wait on. Twelve is a model you can work beside. Fifty is a model that keeps up with you. Same weights, same quant, same software. The only difference is how fast memory can feed the GPU, and the M5 Ultra feeds it four times faster than the M5 Pro and seven times faster than the M6. Also worth noting: 1.2TB/s is roughly 50% more than the M3 Ultra's 819GB/s, which was the previous top of the Apple silicon bandwidth table. The Ultra tier skipped the M4 generation entirely, so this release is two generations of pent-up bandwidth arriving at once. ### Not sure which side of the table you're on? Want to know exactly which hardware your team needs to run local agentic workflows? I consult with enterprise teams on local AI architecture, from sizing the machines to hardening the deployment. Book a Strategy Call ### The M6's dual Neural Engine The M6 is the first Apple chip to ship two 16-core Neural Engines, and Apple says system frameworks can drive both simultaneously for up to twice the peak compute of previous generations. The GPU still owns heavyweight LLM inference, so don't read this as a second lane for your 27B model. Where it gets interesting is agentic plumbing. A working agent stack is never just one model. It is embeddings for vector search, a small classifier routing requests, Whisper transcribing audio, and a vision model reading screenshots, all running beside the main event. Those small concurrent tasks are exactly what a doubled Neural Engine absorbs, keeping the GPU free for the model that needs it. For the always-on appliance pattern from the agent hosting guide, that is a genuinely better node, not just a faster one. ## Super cores, performance cores, and efficiency cores Apple's CPU designs now come in three tiers, and this generation is the first where all three show up in the same buying decision. Super cores are Apple's newest and largest core type, built to win single-threaded work outright. Performance cores carry demanding multithreaded workloads at better power efficiency than a super core would. Efficiency cores absorb background tasks for a fraction of the energy. The M6 is the first base-tier chip to receive super cores, pairing 2 of them with 4 performance cores and 6 efficiency cores, which is how an $899 Mac mini ends up with what Apple calls the world's fastest single-threaded CPU performance. The M5 Ultra takes the opposite shape. Its 36-core CPU is 12 super cores and 24 performance cores with no efficiency cores at all, which tells you exactly what kind of machine it is: one that is always plugged in and always expected to be working. The M5 Pro sits between them with up to 18 CPU cores, enough parallel headroom to keep an inference server, a vector database, and a build pipeline running on one box. Be clear-eyed about what CPU cores do for local AI, though. Token generation lives on the GPU and is fed by memory bandwidth, so no core count rescues a bandwidth-starved chip. Where the CPU tiers matter is everything wrapped around the model: tokenization, the sampling loop, and the orchestration code that runs between agent turns all ride the super cores' single-threaded speed, while efficiency cores are what let an always-on M6 appliance idle at single-digit watts between jobs. Fast plumbing does not make a slow model fast, but slow plumbing can make a fast model feel worse than it is. ## The purchasing decision: one machine or a swarm? Now the fork in the road. All three paths are legitimate, and I have built the swarm side of this in production, so none of the answers here are theoretical. The right one depends on whether your workload is one big model, one mid-size model, or many small ones. ### Price, availability, and the one date that matters Apple opened pre-orders on all of these on August 25, 2026, and they ship on September 22. The M6 Mac mini starts at $899, the M5 Pro Mac mini at $1,699, and the Mac Studio with M5 Ultra at $5,499. There is also a Mac Studio with M5 Max at $2,499 that tops out at 128GB, which deserves a look if you want Studio bandwidth without Ultra money. Then there is the catch, and it is the only scheduling detail in this guide that should change what you actually do. The 512GB configuration is not shipping on September 22. Apple has it arriving in late October. So if that 512GB pool is the whole reason you want the Ultra, and it should be if you are planning to hold a 400B-parameter model in memory, ordering on day one buys you an earlier place in line and nothing else. If your work fits in a smaller configuration, order now and be running in September. ### Scenario A: the M5 Ultra Mac Studio route Buy the Studio if your workload is a single large model that needs to be fast. Zero network latency, plug-and-play setup measured in minutes, and one 512GB pool of 1.2TB/s memory that can hold a 400B-plus parameter model with room left over for a serious context window. The use cases that justify it: complex agentic simulations, high-resolution RAG over large private corpora, and fine-tuning frontier-scale open models on-device, where the data never leaves a machine you own. If your compliance posture requires that last property, the Studio is not a luxury. It is the architecture. ### Scenario B: the single M5 Pro Mac mini route Buy one M5 Pro Mac mini if one mid-size model serves the whole job. A 27B class coding model at 10 to 14 tokens per second on a 64GB machine covers a single developer's local agentic workflow, a team's private RAG service, or an always-on appliance that does one thing well. It is also the configuration to start with when you are not yet sure the workload justifies Ultra money: the box keeps its value as a swarm node or a build machine if you outgrow it. ### Scenario C: the M6 Mac mini cluster route Buy the swarm if you scale horizontally, you are budget-constrained, or your team is experimenting with distributed inference frameworks like Exo or llama.cpp RPC. Five nodes also means five independent failure domains, which is not nothing when the cluster runs production workloads around the clock. Here is the honest math. Five 32GB M6 Mac minis land at roughly $5,000 on Apple's current pricing, and pool 160GB of total RAM. A base M5 Ultra Mac Studio starts at $5,499, so at the entry configuration the swarm and the Studio cost about the same. Push the Studio to 512GB and it climbs well past the swarm, while a single M5 Pro Mac mini at $1,699 costs a fraction of either. The swarm looks great on dollars per gigabyte, until you remember the gigabytes are in five different boxes. To put any of these against what you currently pay a cloud provider, the M4 cluster write-up walks through the cloud spend it replaced and what the hardware cost to run. Factor 5x M6 Mac mini, 32GB Each 1x M5 Ultra Mac Studio, 512GB Approximate cost About $5,000 From $5,499 at base memory, well past the swarm at 512GB Total memory 160GB, split across 5 nodes 512GB, one unified pool Bandwidth to weights 170GB/s per node 1.2TB/s Interconnect Thunderbolt or 10Gb Ethernet, paid on every token UltraFusion, on-package Largest single model Bounded by what layer-splitting and the network tolerate 400B-plus parameters, entirely in memory Many small independent models Excellent, one or two per node Fine, but one GPU is a single queue Failure domains Five One The trap to avoid is treating pooled RAM as if it were unified RAM. When Exo or llama.cpp RPC splits a model's layers across five machines, every generated token has to cross the network between layer groups before the next one can start. Even fast Thunderbolt networking is orders of magnitude slower than memory, so the swarm's 160GB behaves nothing like 160GB in one box. The M5 Ultra's memory is one pool with no hop at all. Where the swarm genuinely wins is the workload it was born for: many independent small models. Whisper on two nodes, embeddings on another, a mid-sized chat model on the rest. That is the pattern our production cluster runs, it is what the setup guide builds, and the M6 is the best node yet made for it. ## Developer tools and frameworks New silicon only matters if software can reach it, and this generation the software kept pace. Apple's frameworks, Core ML and Metal among them, have been updated to tap directly into the M6's Dual 16-core Neural Engine and the GPU Neural Accelerators on both chips. Higher-level runtimes inherit the gains: MLX picks up the accelerators through Metal, so tools like Ollama and llama.cpp benefit without you rewriting anything. One more shift worth flagging for agent builders: developers can now run Apple Foundation Models and App Intents alongside proprietary models on-device. That means the system's built-in model handles the routine intent parsing and summarization while your own open-weight model does the heavy lifting, both on the same box, neither touching a cloud. ## Frequently asked questions ### Is the M5 Ultra, the M5 Pro, or the M6 better for running local LLMs? For a single machine, the M5 Ultra wins decisively. Token generation is bound by memory bandwidth, and the M5 Ultra's 1.2TB/s is roughly four times the M5 Pro's 307GB/s and seven times the M6's 170GB/s. It also holds up to 512GB of unified memory, enough to load models in the hundreds of billions of parameters entirely locally. The M5 Pro Mac mini is the value pick for running one mid-size model well on one quiet box, and the M6 wins on cost per node and idle power, which makes it the better building block for a distributed cluster of smaller specialized models. ### Can a cluster of M6 Mac minis replace one M5 Ultra Mac Studio? Not for a single large model. Five 32GB M6 Mac minis pool 160GB of RAM across a network, but distributed inference frameworks like Exo and llama.cpp RPC pay a network latency penalty on every token, and each node still reads weights at only 170GB/s. One M5 Ultra offers 512GB in a single pool at 1.2TB/s with no network hop at all. A swarm shines when the workload is many independent small models rather than one big one. ### What are the GPU neural accelerators in the M5 Ultra and M6? Both chips embed a Neural Accelerator inside each GPU core, dedicated matrix-multiply hardware in the same place the inference math already runs. Apple cites up to a 30 percent peak AI compute increase on the M6 and up to 4.5x on the M5 Ultra versus their predecessors. The practical effect for LLM work is much faster prompt processing, the compute-bound prefill phase that happens before the first token appears. ### What are super cores, performance cores, and efficiency cores? Apple now splits its CPU designs into three tiers. Super cores chase maximum single-threaded speed, performance cores handle demanding multithreaded work at better power efficiency, and efficiency cores absorb background tasks for minimal energy. The M6 is the first base-tier chip to get super cores, pairing 2 super cores with 4 performance and 6 efficiency cores, while the M5 Ultra scales to 12 super cores and 24 performance cores across its four dies. For local LLMs the GPU still does the heavy lifting, but faster single-threaded CPU work speeds up tokenization, sampling, and the orchestration code wrapped around every agent call. ### How much unified memory does the M6 Mac mini support? The M6 Mac mini configures up to 32GB of unified memory at up to 170GB/s of bandwidth. That is enough for models in the 20 to 27 billion parameter range at 4-bit quantization, such as Qwen3.8-27B, with modest context headroom. ### When does the M5 Ultra Mac Studio ship, and how much does it cost? Apple opened pre-orders on August 25, 2026, and the Mac Studio with M5 Ultra ships on September 22, 2026, starting at $5,499. The Mac Studio with M5 Max starts at $2,499 and tops out at 128GB of unified memory. The 512GB M5 Ultra configuration is the exception: Apple has it arriving in late October rather than on the September 22 date, so a buyer who needs the full memory pool cannot get it at launch. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 Currently Reading ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Current Page Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Bring this talk to your team I speak about zero-cloud AI and on-premise LLM scaling at engineering summits, covering the hardware math, the security model, and the lessons from running Apple silicon in production. For teams that want hands-on depth, the workshop turns this buyer's guide into a working local stack. See AI Speaking Programs Bring the AI Workshop In-House --- ## Page: M6 and M5 Pro Mac mini cluster setup guide **URL:** /projects/ai-mac-cluster/setup-guide **Description:** A step-by-step guide to building a secure M6 or M5 Pro Mac mini cluster with native macOS and PM2. The same blueprint applies to a Mac Studio cluster. # How to build an M6 or M5 Pro Mac mini cluster A comprehensive guide to constructing a high-performance, on-premise cluster for containerized local AI workloads. ## System Architecture Our production architecture runs eight M4 and M4 Pro Mac minis serving requests in a hybrid on-premise and cloud framework. Active workloads are managed through AWS SQS queues. An autoscaler running in Kubernetes coordinates fallbacks, ensuring high availability with zero cloud idle cost. Every step below was proven on that M4 generation, and none of it is specific to it. The rack layout, thermal strategy, network topology, PM2 process management, and cloud fallback all apply unchanged to an M6 or M5 Pro Mac mini cluster. What changes between generations is which box to buy, and that decision is worked through in M5 Ultra vs. M5 Pro vs. M6 for local AI. ## Step 1: hardware selection and network topology Designing a private compute cluster requires choosing the right hardware for the job rather than automatically purchasing the highest-tier systems. A highly successful methodology is to start by benchmarking the base M4 model, then selecting the most cost-effective configuration that fits your specific needs while leaving a comfortable margin of performance headroom. In many distributed server environments, scaling horizontally with multiple inexpensive machines is often far more cost-effective than deploying fewer, highly expensive systems. Additionally, network architecture decisions depend heavily on identifying system bottlenecks. For high-throughput REST API hosting configurations, network input and output capacities are paramount, making 10GbE Ethernet critical. Conversely, for compute-bound systems, such as speech transcription or video processing pipelines, standard 1GbE connections are more than sufficient because the primary processing delay is CPU or GPU execution rather than network transit. By matching your hardware scale and network capacity directly to these system bottlenecks, you establish a lean, highly efficient foundation that avoids over-provisioning. Balancing inexpensive nodes with the appropriate network capacity lets the cluster scale without re-cabling. Once your physical hardware configurations and networking infrastructure are finalized, the next crucial challenge is managing the physical layout and temperature profiles of multiple stacked units. Read on to discover how we architected our rack placement, analyzed real FLIR thermal camera profiles, and kept core temperatures low under continuous peak processing loads. ## Physical layout and thermal management Proper thermodynamic planning is essential when grouping multiple high-performance systems in a dense on-premise space. The aluminum case of each M4 Mac mini serves as a giant physical heat sink, transferring heat away from the internal silicon and dissipating it through the top and sides of the metal chassis. Although Apple's marketing photos typically depict the machines resting flat with their exhaust vents facing down, these compact enclosures do not incorporate a vapor chamber. Consequently, you can position the Mac minis flat or on their side. Both orientations are completely acceptable and thermally stable. For our eight-node cluster, we arranged the Mac minis in a vertical stack layout using wooden risers. We added deliberate vertical separation gaps between each machine in the rack. This vertical layout provides ample physical spacing, allowing cooler ambient air to circulate freely between the aluminum casings and reducing heat transfer from adjacent nodes. Furthermore, we strategically placed our heaviest workload nodes at the top of the stack. These include the M4 Pro systems running our voice transcription pipelines. By positioning them at the peak of the rack, we ensure they have direct contact with rising cooler ambient currents and do not absorb heat radiating from surrounding hardware. The following images are false-color thermal captures that visually illustrate these heating patterns, heat sink performance, and structural airflow benefits across the active rack assembly. Front view showing hottest Mac temp Close-up showing typical Mac temp Side view showing most loaded Mac Show the effect of the cooling riser, 10-20 degrees cooler than case temp These diagnostic images were captured using a specialized FLIR ONE thermal camera. The camera is available for Android platforms and iOS devices, making it an excellent choice for initial hardware assembly and hot-spot inspection during physical setup. While physical thermal imaging is perfect for early validation, you will want to implement continuous software telemetry once your cluster is deployed in production. Utilizing dedicated utilities such as the open-source Mac Stats tool allows you to pull live thermal values directly from internal macOS sensors. For scaling setups, you can configure telemetry scripts to forward these temperature readings directly into central dashboards like Grafana or Amazon CloudWatch to monitor the thermal health of every node in real time. Real-time dashboard visualization representing remote CPU core thermal loads and system health status. ## Step 2: base operating system configuration To optimize performance and minimize idle background overhead, we run native macOS across our physical nodes. Running a full graphical user interface introduces an idle memory overhead of a few hundred megabytes, but the benefit of accessing the complete, uncompromised compute resources of the M4 silicon makes this minimal memory tax well worth paying. We begin by stripping out unnecessary consumer software, including the default Chess application, to minimize the overall attack surface. Furthermore, to maximize efficiency and eliminate background memory consumption, we disable Wi-Fi and Bluetooth entirely, routing all cluster communications through our high-speed wired Ethernet network. To guarantee enterprise-grade security at the local hardware level, every node is configured with FileVault disk encryption enabled and the native macOS Firewall fully active. Maintaining high availability and uninterrupted remote control requires adjusting several system preferences. We apply the following specific macOS configuration parameters to each node: - Energy Mode: High Power is selected to ensure the CPU and GPU operate at peak potential. - Prevent automatic sleeping when the display is off: Enabled to keep system threads constantly active. - Put hard disks to sleep when possible: Disabled to prevent any drive spin-down latency. - Wake for network access: Enabled to support remote wake signals. - Start up automatically after power failure: Enabled to recover cluster operations automatically when utility power is restored. - Turn off display when inactive: Set to Never to ensure background rendering and process loops are not interrupted. - Log in automatically: Enabled to bypass interactive login prompts during system restarts. - Open at Login: Configured with a custom Automator application that executes the required initialization command line scripts. - Screen Sharing and Remote Login: Enabled to facilitate remote cluster administration and secure shell commands from a primary laptop. - Display Simulation: Active HDMI dummy plugs are inserted into each node to simulate a physical monitor. Without these dummy plugs, the graphics subsystem can sleep, causing a black hanging screen when initiating Screen Sharing remote access sessions. Once these core OS optimizations and security baselines are established, we capture a backup of the finalized clean system state using Time Machine. This base image enables our team to quickly restore or deploy new nodes to an identical, standardized starting point. ## Step 3: establishing secure remote access Managing an on-premise compute stack requires reliable remote administration capabilities. However, exposing default remote access protocols such as VNC, Screen Sharing, or secure shell access directly to the public internet is wildly insecure and leaves nodes vulnerable to automated brute-force attacks. To safeguard the cluster, we configured a private Virtual Private Network topology rather than mapping ports publicly. Private overlay network solutions allow for encrypted communication between devices without configuring complex router port forwarding or exposing local firewall ports. Two popular and reliable tools for this tool are ZeroTier and Tailscale. Both services enable you to group physical cluster nodes and administration laptops into a single, secure virtual local area network. For our cluster, I personally use ZeroTier but have heard good feedback about Tailscale too. ## Step 4: determining the right process management solution When orchestrating physical node clusters, it is common to deploy container orchestration platforms such as k3s. However, container runtimes and many Linux distributions do not natively support or expose core Apple silicon hardware features, including the unified memory GPU, Metal compilation, or the Apple Neural Engine. Because we host heavy AI workloads, we chose to bypass these container orchestration systems to prevent losing critical GPU and Neural Engine hardware acceleration. Rather than managing containers, we monitor node processes directly using the lightweight process manager pm2 and sync all system logs and telemetry to Grafana. The machines run all operational servers and processing pipelines under a restricted, non-admin system user named web, establishing a hard boundary against unauthorized lateral movement and keeping application environments jailed in compliance with strict boundary defense guidelines. ## Step 5: designing hybrid cloud fallbacks via AWS SQS To prevent service interruption if physical power or external network connectivity drops, route all incoming workload request payloads into an AWS SQS queue. The local cluster nodes continuously pull from this queue to perform processing. Concurrently, configure a cloud-based autoscaler to monitor the queue depth. If the queue length increases beyond five messages or the local cluster goes offline, temporary cloud instances automatically spin up to handle the load. Once the local physical cluster resumes operation or the load decreases, the cloud workers scale back to zero replicas. This hybrid model keeps cloud idle costs at zero under normal conditions. ## On-premise tinkering versus data center colocation While deploying physical Mac minis in a remote collocated data center cabinet provides redundant power loops and high bandwidth pipes, starting with an on-premises setup is ideal for development. Having physical access lets me tinker iteratively, swap network components, and adjust racks easily. I have decided to move to a collocated data center once I have stopped tinkering for a few weeks on the unit. ## Operating system selection: native macOS versus bare-metal Linux A common point of discussion is whether to flash these physical nodes to run bare-metal Linux distributions. While headless Linux consumes slightly less idle operating system memory than macOS, native macOS remains the definitive choice for newer Apple silicon generations. Native Apple Neural Engine acceleration and Metal GPU compilation drivers are not fully supported or stable on Linux for newer M3 and M4 series hardware. Attempting to run bare-metal Linux results in losing the unified memory GPU acceleration libraries that make Apple silicon exceptionally efficient for local large language model inference. Sticking with macOS ensures full AppleCare warranty compliance and clean system updates. ## Mitigating non-ECC memory risks in consumer hardware Enterprise server architects often point out that the Mac mini uses standard consumer memory instead of Error correcting code memory. In standard database or operating system servers, memory bit flips can cause silent data corruption or sudden crashes, which makes Error correcting code memory desirable. For our cluster, we mitigate this risk through our architectural model. The Mac minis function as stateless, ephemeral workers pulling from a queue. If a physical node encounters a memory anomaly and crashes, the task is not lost. The message simply times out in the queue, and our cloud failover instance or another cluster node automatically picks it up for retry. This stateless architecture makes consumer-grade hardware perfectly safe and highly cost-effective for high-frequency AI inference tasks. ## Power footprint and cost efficiency Another common critique of localized compute clusters is the high cost of electricity and air conditioning. Traditional server hardware requires dedicated high-amp circuits, custom server rooms, and massive cooling units that quickly offset any software licensing savings. The M4 Apple silicon platform bypasses this entire infrastructure burden. Because each base M4 mini draws only four watts when idle and under thirty-two watts under peak GPU inference load, our entire eight-node cluster operates at under three hundred watts under peak load. This power draw is less than a standard consumer gaming desktop or a simple space heater. Consequently, standard office electrical outlets and routine in-window cooling are more than sufficient to maintain stable core temperatures of seventy-two degrees Celsius under heavy continuous loads. To safeguard our cluster against sudden power drops or local outages, we integrate 510W CyberPower battery backups. We route three Mac mini systems per backup unit to maintain a comfortable power margin. Because each M4 Pro model exhibits a peak power draw of 140 watts, loading three active nodes per unit ensures a combined full-blast power draw of 420 watts, leaving ample headroom under the 510-watt maximum limit. When installing these backups, be careful to connect the systems exclusively to the battery-backed outlets, as the left side of these units only provides surge protection without battery fallback. To prove how incredibly efficient these systems are in practice, we tracked the real production runtime consumption of one of our M4 Pro speech transcription nodes over a complete week. Because Apple silicon is highly optimized for deep sleep transitions and low-power processing, the transcription node consumed a total of only 915 watt-hours during a full seven days of continuous operation. Real-world production power tracker showing a total consumption of 915 watt-hours over a full week of active uptime. Extrapolating this weekly consumption over a full year of fifty-two weeks results in an annual energy requirement of only 47.58 kilowatt-hours per node. This makes physical local hosting remarkably inexpensive, even in regions with high residential energy rates. For example, our cluster operates in a region serviced by San Diego Gas and Electric, where residential utility rates are notoriously high. During a recent billing cycle, our residential electricity charges totaled $344.42 for 754 kilowatt-hours of usage, which translates to a high rate of approximately $0.46 per kilowatt-hour. We calculated this by subtracting the irrelevant $37.16 natural gas service charge from the total utility statement of $381.58. Local utility statement reflecting high residential electricity rates. Even when subjected to these high utility rates, the annual cost to power one of our production speech transcription nodes is only $21.89. This extremely low operational cost makes localized Apple silicon clusters exceptionally cost-effective compared to cloud serving models. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 Currently Reading ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Current Page Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Want this blueprint live, on your stage? I keynote engineering conferences on local AI infrastructure and run hands-on workshops that take teams from hardware selection to a hardened deployment. Bring the session to your event or your engineering organization. See AI Speaking Programs Bring the AI Workshop In-House --- ## Page: Qwen 3.8 on a Mac: local agentic coding on Apple silicon **URL:** /projects/ai-mac-cluster/agentic-coding **Description:** Run Qwen 3.8 27B locally on Apple silicon with Ollama and Zoo Code. Unified memory needs, context window setup, and the bandwidth numbers that decide it. # Run Qwen 3.8 on Apple silicon, without rate limits Qwen3.8-27B runs locally on a Mac Studio with Ollama and Zoo Code. Here is the setup, the unified memory you actually need, and the honest performance ceiling. As much as I love Mac minis, this is the one job where you want a Studio. An Ultra-class Studio, to be specific. Read on for why. Photo by Joey Banks on Unsplash. ## The Rate-Limit Frustration You're in the flow. The agent is mid-refactor, the tests are green, and you're about to ship. Then the spinner appears. Rate limited. It happens with Claude. It happens with OpenAI. It happens with Antigravity. You hit the hourly token ceiling, the daily request cap, or the concurrent-session limit, and your momentum dies. You wait. You switch accounts. You downgrade your prompt. You lose the thread. A bigger API quota just moves the ceiling further out. What actually fixes this is a local model you own, running on hardware you control, available around the clock with no meter running. The catch is that the hardware has to be right, and that is the part most guides gloss over. This one doesn't. ## Which Qwen 3.8 is this? Worth settling before anything else, because the naming trips people up. Three different things get called some version of Qwen 3.8, and only one of them is the subject of this guide. Name What It Actually Is Qwen3.8-27B The open-weight 27-billion parameter model released under Apache 2.0. This is the one you can download and run on your own Mac, and the one this guide covers. Qwen 3.8-Max A much larger flagship in the same generation, available through an API rather than as weights you can download. You cannot run this one locally. Qwen3-8B An older 8-billion parameter model from the previous Qwen3 generation. The similar spelling is a coincidence of version numbering, not a smaller edition of Qwen 3.8. Everywhere below, Qwen 3.8 means Qwen3.8-27B. On Ollama it is simply the qwen3.8 tag, which is where the spelling without a space comes from. And if you're still weighing Qwen against the other open-weight contenders, the local coding model comparison covers the whole field. This page assumes you've settled on Qwen and want it running. ## Qwen 3.8 hardware requirements on Apple silicon Start here, because this is the step that decides whether the rest of the guide works. The model this article recommends is Qwen3.8, a 27-billion parameter model with a 256K context window. The default tag Ollama pulls is a 4-bit quant that lands at 18GB on disk, and it has to sit in unified memory to run at speed. That single number sets the floor. A 16GB Mac cannot hold this model. 32GB runs it but leaves little headroom for context. 64GB is where local agentic coding stops feeling like a compromise, for a reason covered in the context window step below: macOS only lets the GPU address part of unified memory, so the memory Ollama sees is always smaller than the number on the spec sheet. Unified Memory What You Get 16GB Will not run the 18GB model. Pick a smaller model or a bigger Mac. 24GB to 32GB Runs, but the context window is the constraint. The GPU sees roughly two-thirds of unified memory, so Ollama defaults to a 4k context here. Workable for single-file edits and short tasks. 48GB Comfortable for the weights, but the GPU sees around 36GiB, which lands in Ollama's 32k default tier rather than the full 256k. 64GB The sweet spot. Roughly 48GiB goes to the GPU, which is where Ollama's 256k default tier begins, with room for a long agent context on a real codebase. 96GB and up Headroom for larger models or several loaded at once. Spend here only if bandwidth is already high. Capacity decides whether the model runs at all. Memory bandwidth decides how fast it feels. Those are two different specs, and the Under the Hood section has the full table of Apple silicon bandwidth figures so you can size the second one properly. ## How to run Qwen 3.8 on a Mac with Ollama Seven steps. No cloud account, no API key, no meter. Budget most of the wall-clock time for the 18GB model download. ### 1. Get Ollama Ollama is the local runtime that serves open-weight models over a clean HTTP API. Grab the macOS build from the Ollama download page. ### 2. Pull Qwen3.8 Qwen3.8 is the model I recommend for local agentic coding. It is strong at multi-step reasoning and tool use, it follows agent skills and markdown instruction files closely enough for agent mode to work, and thinking mode is on by default with the reasoning depth tunable per request. Browse the model card on the Ollama library, then pull it from your terminal: $ ollama pull qwen3.8 That pulls qwen3.8:latest, an 18GB 4-bit quant. If you have memory to spare and want higher fidelity, qwen3.8:27b-q8_0 is 30GB and qwen3.8:27b-bf16 is 56GB. One thing worth knowing before you settle on Ollama for this: I benchmarked it against Apple's own MLX runtime on the same weights, and on a dense model at 4-bit MLX generates about 40 percent more tokens per second. The advantage disappears at 8-bit and reverses on mixture-of-experts models, so it is not a straight upgrade. The full measurements are here, along with the prompt-caching trap that makes most published comparisons wrong. ### 3. Raise the context window This is the step that is easy to skip, and it is a common reason people conclude that local agents are useless. Ollama picks a default context length based on available VRAM: under 24GiB you get 4k tokens, 24GiB to 48GiB gets you 32k, and 48GiB or more gets you the full 256k. Those thresholds are VRAM, not unified memory, and on Apple silicon the two are not the same number. macOS reserves part of unified memory for the system and hands the GPU roughly two-thirds of it on machines up to 36GB and roughly three-quarters above that. A 48GB Mac therefore presents about 36GiB to Ollama and lands in the 32k tier, not the 256k one. That is the real reason the sweet spot is 64GB rather than 48GB. A 4k context is not an agent. It is an autocomplete that forgets the file it just opened. Ollama's own context length documentation is explicit about it: tasks that need large context, including agents and coding tools, should be set to at least 64000 tokens. Treat that as the floor. For agentic coding, give it the model's full window: $ OLLAMA_CONTEXT_LENGTH=262144 ollama serve 262144 is the 256k window Qwen3.8 was trained for, and it is the same value Ollama picks on its own once it sees 48GiB of VRAM. Setting it explicitly means you get that window regardless of how Ollama reads your hardware. On a 64GB Mac or larger there is room for it. On a smaller machine the setting will still apply, but the weights plus a context that size will not fit in what the GPU can address, and Ollama will spill to the CPU and slow to a crawl. Drop to 65536 there. The Ollama desktop app exposes the same setting as a slider in its settings menu. Once the model is loaded, ollama ps confirms both the context window it was given and whether it is actually running on the GPU. This is also the honest reason the memory tiers above matter. A longer context costs memory on top of the 18GB of weights, so context length and unified memory are the same budget spent twice. ### 4. Install VS Code If you don't already have it, grab VS Code. ### 5. Install the Zoo Code extension Add the Zoo Code extension from the VS Code marketplace. It gives you a set of agent modes in the editor, including Architect for planning, Debug for diagnosis, and Code for everyday edits and file operations. ### 6. Connect to Ollama In Zoo Code's settings, select ollama as the API provider, leave the base URL at the local default of http://localhost:11434, and enter qwen3.8 as the model. The Zoo Code Ollama provider documentation covers the full option list. One detail worth knowing: Zoo Code defers to the model's num_ctx as Ollama reports it rather than setting a context length of its own. That is why step 3 comes before this one. If you raise the context after the model is already loaded, restart the Ollama server so the new value takes effect. ### 7. Start coding Put the tool in Code mode and start working. It picks up your agent skills and markdown instruction files the same way a cloud coding agent does. No rate limit. No meter. No waiting for a quota to reset. That's one developer off the meter. Rolling the same stack out to a whole engineering team is a different project, and the first question is not which model to run. It's whether your codebase is something an agent can actually work in. I keep a ten-point agent-ready codebase audit for exactly that, and it takes about an hour to run against a real repo. If it comes back more red than green, closing that gap is what my hands-on agentic coding workshop is built around, from first local model to a production setup the team actually uses. ## Under the Hood If you're choosing hardware or wondering why one Mac feels faster than another, these are the four things that actually matter for local agentic coding. ### 1. Prompt processing is bound by GPU speed The first pass over your prompt is compute-bound. The GPU has to process every token in the context window before it can start generating. Newer GPU architectures with more cores and wider memory interfaces finish this pass faster. An M3 Ultra will chew through a long prompt noticeably quicker than an M1, even if the two chips end up generating tokens at a similar clip. This is the spec you feel most in agentic coding specifically, because an agent re-reads a large context on every turn. It is the difference between a pause you ignore and a pause you alt-tab away from. ### 2. Token generation is bound by memory bandwidth Once the model starts writing, every token it generates requires reading the model's weights back out of memory. That makes bandwidth the ceiling. Raw compute barely enters into it. The practical upshot is that bandwidth tiers matter more than product names. An M1 Ultra at 800GB/s and an M3 Ultra at 819GB/s generate tokens at roughly the same speed despite two generations between them, because they sit in the same bandwidth class. A recent Max chip lands close behind. A base chip is a different category of experience entirely, and no amount of extra RAM changes that. Worth stating plainly, since it causes a lot of confusion: there is no M4 Ultra. Apple's Ultra tier skipped the M4 generation entirely, and the drought ended with the M5 Ultra, which now tops the bandwidth chart at 1.2TB/s, followed by the M3 Ultra, the M5 Max, and the M4 Max. This is the whole case for the Studio, and I say that as someone who runs a rack of Mac minis. The mini tops out at the M5 Pro and 307GB/s. An M5 Ultra Studio reads the same weights at 1.2TB/s. Identical model, identical quant, four times the memory bandwidth feeding the GPU. Here's the Apple silicon lineup, sorted by generation: Chip Memory Bandwidth Class M1 68.25 GB/s† Base M1 Pro 200 GB/s Pro M1 Max 400 GB/s Max M1 Ultra 800 GB/s Ultra M2 100 GB/s Base M2 Pro 200 GB/s Pro M2 Max 400 GB/s Max M2 Ultra 800 GB/s Ultra M3 100 GB/s Base M3 Pro 150 GB/s Pro M3 Max 300–400 GB/s Max M3 Ultra 819 GB/s Ultra M4 120 GB/s Base M4 Pro 273 GB/s Pro M4 Max 410–546 GB/s Max M5 153 GB/s Base M5 Pro 307 GB/s Pro M5 Max 460–614 GB/s Max M5 Ultra 1.2 TB/s Ultra M6 170 GB/s Base † Apple has never published a bandwidth figure for the original M1. The value is the widely reported number derived from its memory interface width and speed, and should be treated as a close approximation rather than a vendor specification. Every other figure in this table comes from Apple. #### Sources - Apple MacBook Pro specs — M5 at 153 GB/s, M5 Pro at 307 GB/s, M5 Max at 460 GB/s and 614 GB/s - Apple Mac Studio specs — M3 Ultra at 819 GB/s, M4 Max at 410 GB/s and 546 GB/s - Apple Mac mini specs — M4 at 120 GB/s, M4 Pro at 273 GB/s - Apple M6 and M5 Ultra announcement — M5 Ultra at 1.2 TB/s, M6 at 170 GB/s - Apple MacBook Pro 14-inch M3 Pro and M3 Max tech specs — M3 Pro at 150 GB/s, M3 Max at 300 GB/s and 400 GB/s - Apple MacBook Pro 14-inch M3 tech specs — M3 at 100 GB/s - Apple Newsroom: M2 Pro and M2 Max — M2 Pro at 200 GB/s, M2 Max at 400 GB/s - Apple Newsroom: M2 — M2 at 100 GB/s - Apple Newsroom: M2 Ultra — M2 Ultra at 800 GB/s, stated as twice that of M2 Max - Apple Newsroom: M1 Ultra — M1 Ultra at 800 GB/s - Apple Newsroom: M1 Pro and M1 Max — M1 Pro at 200 GB/s, M1 Max at 400 GB/s ### 3. Memory capacity is a floor, not a dial You need enough unified memory to hold the weights and the context window at the same time. If it doesn't fit, it doesn't run, or it spills and crawls. That's the hard floor, and for Qwen3.8 the floor is 18GB of weights before a single token of your codebase is loaded. Above the floor, returns diminish quickly. Going from 32GB to 64GB buys real context headroom and is the upgrade most people should make. Going from 64GB to 128GB buys very little for this model, because the weights and a generous context already fit. At that point the money is better spent on bandwidth, which is the spec that never stops mattering. ### 4. One task at a time, around the clock An agentic coding workload saturates the GPU. You're not going to run two heavy coding sessions in parallel on a single Mac and expect both to stay fast. Plan for one task at a time. But here's the trade-off that makes it worth it: you can run that one task continuously without a rate limit. No hourly cap. No daily ceiling. No concurrent-session limit. The machine doesn't care if it's 3 AM. That should soften the blow of single-task throughput. It also points at the only real way to scale this, which is more machines rather than a bigger one. That's the entire argument behind the Mac cluster build documented in the rest of this series. ## Qwen 3.8 coding performance: what to expect A 27-billion parameter model running at 4-bit precision on a desktop is not a frontier model. Here is where the line actually falls. Qwen publishes benchmark results for Qwen3.8-27B on its model card, including 73.0 on Terminal-Bench 2.1 and 61.7 on SWE-bench Pro. Those are vendor-reported numbers on evaluations Qwen selected and in some cases modified, so treat them as a directional signal rather than an independent verdict. They are consistent with what the model feels like in practice, which is the useful part. It handles the work that fills most of a day: implementing a well-specified function, writing tests, tracing a bug through a few files, refactoring a module, and writing documentation. It follows instructions in markdown files consistently enough to stay on task, which is what makes agent mode viable at all. It struggles where the frontier models still earn their price: sprawling multi-file architectural changes, subtle reasoning about unfamiliar library internals, and long autonomous runs where a small early mistake compounds. Thinking mode helps and costs tokens, so on lower-bandwidth hardware you feel that trade directly. Think of it as triage. The local model takes the volume work with no meter running. The cloud model gets the two or three genuinely hard problems you hit in a day. That split is also what makes the rate-limit ceiling stop mattering, because you stop spending your quota on boilerplate. Getting one developer onto that split takes an afternoon. Getting a whole engineering org onto it is a people problem, and it is the one I spend most of my time on. If that is the problem sitting on your desk, my AI speaking programs cover how to make the triage stick across a team, fees and formats included. ## Frequently asked questions ### Is Qwen 3.8 the same as Qwen3-8B? No. Qwen3-8B is an older 8-billion parameter model from the Qwen3 generation. Qwen 3.8 is a newer generation, and the open-weight model in it is Qwen3.8-27B, a 27-billion parameter model released under Apache 2.0. There is also Qwen 3.8-Max, a much larger API-only flagship in the same generation. This guide covers Qwen3.8-27B, the one you can actually run on a Mac. ### Can you run Qwen 3.8 on a Mac? Yes. Qwen3.8-27B runs on Apple silicon through Ollama. The default 4-bit build downloads at 18GB and has to fit in unified memory, so 32GB is the practical floor and 64GB is where it stops feeling like a compromise. There are also mlx tags that target Apple’s own machine learning framework. ### How much RAM do you need for local agentic coding on a Mac? The default Qwen3.8 build is a 27-billion parameter model that downloads at 18GB. A 16GB Mac cannot hold it. 32GB is the practical floor, and 64GB is what unlocks Ollama’s largest default context window, which is what agentic coding actually consumes. The gap exists because macOS gives the GPU only about two-thirds to three-quarters of unified memory, so a 48GB Mac presents roughly 36GiB to Ollama. ### Do you need a Mac Studio with an Ultra chip, or will a Mac mini work? Token generation speed tracks memory bandwidth. The M5 Ultra Mac Studio at 1.2TB/s is the new ceiling, and an M3 Ultra at 819GB/s or an M5 Max at 614GB/s remain comfortable. An M4 Pro Mac mini at 273GB/s runs the model but generates roughly a third as fast as the M3 Ultra. There is no M4 Ultra, so the Ultra tier means a Mac Studio with the M5 Ultra, the M3 Ultra, or an older M1 or M2 Ultra. ### Why is my local coding agent forgetting context or failing mid-task? Ollama sizes its default context window by available VRAM, and under 24GiB that default is only 4k tokens. Ollama’s own documentation puts the floor for agents and coding tools at 64000 tokens, but for agentic coding set OLLAMA_CONTEXT_LENGTH=262144 to get the full 256k window Qwen3.8 was trained for. That needs a 64GB Mac or larger; on a smaller machine drop to 65536 so the weights and context still fit in memory the GPU can address. Set it before starting the agent, not after. ### Can you run more than one agentic coding session on a single Mac? An agentic coding workload saturates the GPU, so plan on one heavy session per machine. Scaling out means adding machines, which is the argument for a Mac mini or Mac Studio cluster rather than a single larger Mac. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 Currently Reading ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Current Page Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Ready to code without limits? I keynote conferences on zero-cloud AI, and my hands-on workshop, Modernize Your Engineering Org for the AI Era, takes engineering teams from their first local model to a production agentic coding stack. See AI Speaking Programs Book the Agentic Coding Workshop --- ## Page: Private AI agent hosting on M6 & M5 Pro Mac minis **URL:** /projects/ai-mac-cluster/agent-hosting **Description:** Configure an M6 or M5 Pro Mac mini cluster as an always-on, low-power private AI host for local LLMs, Whisper transcription, and vision models. # Local AI agent hosting on M6 and M5 Pro Mac minis Turn a compact, 4-watt idle Mac mini into a powerful, secure, 24/7 private host for autonomous agentic workflows. ## The era of private AI appliances As organizations adopt autonomous agentic workflows, relying on public cloud APIs introduces critical risks regarding data leakage, high operational latency, and runaway token expenses. An M6 or M5 Pro Mac mini acts as a private, self-contained AI appliance that processes proprietary documents and database integrations locally within your network boundaries. ## Why M6 and M5 Apple silicon is ideal for AI agents Unlike transient batch workloads, autonomous agents require reliable, continuous, and efficient system resources. The M6 and M5 system-on-chip platforms deliver outstanding capabilities specifically designed for local background processing: - Supreme Idle Efficiency: The base M6 Mac mini, built on Apple's first 2nm process, draws just a few watts of power when idling, allowing you to run background processes continuously without impacting electricity bills or thermal wear. - Unified Memory Advantage: The shared UMA architecture allows agent frameworks to rapidly call local LLM endpoints, processing massive token context pools with zero physical data transfer bus delays. - Neural Engine Acceleration: The M6's Dual 16-core Neural Engine handles light embedding tasks, vector searches, and system routines at maximum processing speed while freeing the GPU for primary inference loops. ## Selecting your local AI use case When deploying a local Apple silicon cluster, the primary design challenge is aligning your physical computing power with the specific software architecture of your choice. Depending on your workload requirements, you can configure your cluster to handle one of four primary local AI use cases: ### 1. Serving local large language models For standard text completion, retrieval-augmented generation, and interactive chat, you can serve open-weight models using lightweight local runtime engines. I recommend using Ollama for dependable headless daemon hosting or LM Studio for comprehensive local testing and API endpoint hosting. By assigning a custom Domain Name System hostname to your physical node's unique IP address within your ZeroTier private network, you establish a secure, encrypted link back to your hardware cluster from anywhere in the world without exposing your endpoints to the public internet. To interact with these models, you can connect the local API endpoints to clean front-end interfaces. For web browsers, OpenWebUI provides a feature-rich, self-hosted web chat interface. For mobile devices, Invoke serves as a native iOS app that connects directly to your private local API endpoints, speaking to both Ollama and LM Studio over HTTP. ### 2. Autonomous agents If you want to run fully autonomous workflows, you can use advanced agent orchestration frameworks like Hermes or OpenClaw. These libraries enable AI agents to execute multi-step plans, call external APIs, and run arbitrary terminal commands to solve complex problems. I am too nervous to run fully autonomous agent frameworks on our primary cluster because of the severe security risks associated with a lack of a blast radius. Letting an AI model execute shell commands on your local system with write access to the filesystem is highly risky. If the model goes off course or is subjected to prompt injection, it could accidentally delete directories, leak secrets, or compromise the host machine. From a compliance perspective, running un-sandboxed autonomous agents on local production hosts is highly problematic for SOC 2 audits, as it violates basic data isolation and lateral movement prevention principles. If you choose to deploy these tools, you must execute them inside strictly jailed virtual environments or sandboxed containers to limit their execution boundary. ### 3. High-frequency speech transcription Speech-to-text processing is one of the most cost-effective workloads to repatriate to Apple silicon. Instead of paying continuous per-minute API fees to public clouds, you can run highly optimized speech transcription pipelines locally. We use whisper.cpp to perform high-speed, local speech transcription. The C/C++ port of OpenAI's Whisper model compiles natively on macOS, utilizing Apple silicon Unified Memory Architecture and Metal GPU shaders to transcribe multi-hour audio recordings in minutes with zero external network dependencies. ### 4. Local vision and multimodal workloads Processing images, performing optical character recognition, and running visual reasoning tasks can be handled entirely on local hardware. Vision-language models, often referred to as VLMs, have evolved to operate exactly like standard text large language models on Apple silicon. Multimodal models like Llama 3.2 Vision and Qwen 2.5 VL can be compiled and run locally using the same Ollama or LM Studio backends. By utilizing unified memory, the GPU can load both visual and textual weights into the same memory space, enabling instant image analysis, document scanning, and automated UI inspection without any data leaving your local host. ## Predictable workloads and cloud repatriation costs For small and medium enterprises, running continuous compute pipelines under elastic cloud APIs is highly cost prohibitive. While the cloud is excellent for global elasticity and highly variable traffic peaks, predictable everyday workloads belong on owned local hardware. This is not just a small-company calculation. Netflix's AI platform team wrote about serving LLMs in-house rather than routing every request through external APIs, folding open-weight and custom models served on vLLM into the same production scoring infrastructure that already handles the rest of their model traffic. Their reasoning maps directly onto a Mac mini rack: operational fit beats raw benchmark performance, and owning the serving path is what makes fast iteration on your own workloads possible. For example, our voice transcription workloads were previously running on Google Cloud's Speech-to-Text API, which costs $0.016 per minute. Migrating this work to Whisper models running locally on our Mac mini cluster yielded immediate savings. An M5 Pro Mac mini needs only 1 GPU core and 2 GB of RAM to keep up with a real-time speech-to-text transcript. Consequently, each 64 GB Mac mini node can run between 10 to 20 parallel transcription pipelines, keeping up with real-time stream processing with zero variable usage fees. Repatriating these workloads provides substantial long-term cost benefits and guarantees complete data residency control over sensitive files. ## Understanding the scale and memory constraints of local agents When planning a private AI deployment, it is vital to match the target workload to the appropriate system memory bandwidth. A cluster of physical M6 or M5 Pro Mac minis is highly optimized for hosting specialized small-to-medium language models that handle parallel asynchronous tasks such as voice transcription, database querying, and vector database generation. Heavy autonomous software engineering agents are a different workload. An agent that parses an entire code repository re-reads a very large context on every turn, so it is constrained by unified memory bandwidth and capacity in a way that a transcription pipeline never is. An M5 Pro Mac mini at 307 GB/s runs a mid-sized coding model, but it does not feel like the hosted tools you are used to. That is the point where the hardware conversation shifts from Mac mini to Mac Studio. The M5 Ultra Mac Studio moves that ceiling to 1.2 TB/s of memory bandwidth and up to 512 GB of unified memory in a single pool, enough to load models in the hundreds of billions of parameters entirely locally. I cover the model selection, the context window configuration, and the full Apple silicon bandwidth table in local agentic coding on Apple silicon, and the cluster-versus-monolith purchasing decision in M5 Ultra vs. M6 for local AI. The Mac mini cluster remains the workhorse for high-frequency specialized micro-agent services, and a Mac Studio is the machine you add when you want the coding agent too. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 Currently Reading ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Current Page Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Ready to build private AI agents? I speak about private AI appliances and agent hosting at engineering events, and my hands-on workshop shows teams how to configure local agent orchestrators, private RAG architectures, and hardened deployments. See AI Speaking Programs Bring the AI Workshop In-House --- ## Page: Local AI use cases: local vs. cloud AI architecture **URL:** /projects/ai-mac-cluster/local-ai-use-cases **Description:** Where M5 Ultra Mac Studios and M6 Mac mini swarms beat cloud APIs on cost, latency, and privacy, plus the hybrid local-first framework. # The enterprise guide to local vs. cloud AI architecture Where the M5 Ultra and M6 Mac mini swarms outperform cloud APIs on cost, latency, data privacy, and physical AI, and when a hybrid approach wins. Once the hardware is on your desk, every additional token is free. That one fact reshapes more architecture decisions than any benchmark. ## The architectural premise Every AI infrastructure conversation I have with engineering leaders eventually lands on the same question: which of our workloads actually need the cloud? Not which ones happen to run there today. Which ones need it. The honest answer is fewer than the invoice suggests. Local Apple silicon offers three properties no API vendor can sell you: zero-marginal-cost token generation, an air gap between your data and the internet, and compute that physically sits next to the sensors and actuators that need it. This guide maps the five workload categories where those properties dominate, the hardware that fits each one, and the decision matrix for everything in between. It is the strategy layer on top of the production cluster this series documents. ## Air-gapped privacy and regulatory compliance Some data should never transit a network you don't own. Legal teams auditing contracts, healthcare organizations processing PHI under HIPAA and GDPR, finance groups running internal models, and any company analyzing its own proprietary IP all share the same constraint: the value of the analysis is capped by the risk of the transmission. Local inference deletes that risk rather than managing it. There is no data processing agreement to negotiate with a machine in your own server closet. - Representative workloads: contract review and clause extraction, PHI de-identification and summarization, internal financial modeling, and RAG over proprietary research that competitors would pay dearly to see. - The hardware: an M5 Ultra Mac Studio with 512GB of unified memory. At 1.2TB/s of bandwidth it holds a frontier-scale open-weight model and your full enterprise RAG context in one memory pool, on one machine, behind your own badge reader. - The cloud contrast: no data transmission, no vendor logging, no retention policy to audit, and no clause about model training buried in an updated terms of service. The compliance article in this series covers how this architecture satisfies ISO 27001:2022 and SOC 2 auditors. The ROI trigger here is not a spreadsheet. It is the first time your general counsel asks where the embeddings live. ## High-volume agentic coding and development environments Agentic coding is the workload that breaks cloud pricing models. An autonomous agent re-reads its context on every turn, runs in loops, and works around the clock. Meter that by the token and the bill curve bends the wrong way. Meter it by the kilowatt and it flattens. - Representative workloads: local IDE autocomplete, multi-file codebase indexing, automated test-case generation, and recursive CI/CD loops that generate and validate synthetic data all night. - The hardware: a swarm of four M6 Mac minis on Thunderbolt 5, orchestrated with Exo. Each node serves its own model, so four agents run in parallel without queueing behind one GPU. - The cloud contrast: zero API rate limits, zero token billing during continuous multi-agent loops, and first-token latency measured on a local bus instead of a WAN. An agent that never waits on a rate limiter is a different product than one that does. The agentic coding article works through the model selection and the memory bandwidth math, and the M5 Ultra vs. M6 buyer's guide settles which box to buy for it. ### Transitioning your team to local AI infrastructure? I help engineering leaders design air-gapped local AI deployments, evaluate M5 Ultra and M6 hardware ROI against their current cloud spend, and ship private agentic workflows their auditors sign off on. Book an Enterprise Architecture Review ## Physical AI, robotics, and edge computing A chatbot can tolerate a 400ms round trip to a data center. A robot arm cannot. When the model output moves metal, the network hop stops being a cost problem and becomes a safety problem, and the only fix is to put the compute on the machine. - Representative workloads: real-time ROS 2 sensor fusion across LiDAR and spatial cameras, vision-language-action model execution with OpenVLA or Octo, and physical kinematics simulation before a motion plan ever runs on hardware. - The hardware: an onboard M6 Mac mini. The 2nm process keeps the power draw inside a mobile battery budget, and the Dual 16-core Neural Engine absorbs the perception models while the GPU runs the VLA policy, keeping actuator control latency local and deterministic. - The cloud contrast: no round-trip latency in the control loop and no internet dependency during autonomous navigation. A robot that stops working when the Wi-Fi does is a demo, not a product. ### Building autonomous physical AI systems? I run hands-on workshops for R&D teams on deploying agentic pipelines and edge inference on Apple silicon, from swarm orchestration to the latency budget that keeps a control loop safe. Request a Team Workshop Syllabus ## High-density batch execution and model training Some jobs don't need to be fast. They need to be relentless. Batch workloads run for twelve hours at full utilization, which is exactly the shape of job that makes per-token pricing look like a rounding error at first and a line item with its own budget review by Q3. - Representative workloads: overnight log analysis, terabyte-scale document processing, domain-specific synthetic dataset creation, and adversarial red-team evaluation of your own models before an attacker volunteers to do it for you. - The hardware: multi-node Apple silicon clusters running llama.cpp RPC or MLX swarm topologies. Batch work parallelizes cleanly across nodes because each document is independent, which is the workload swarms were born for. - The cloud contrast: unlimited continuous execution with no exponential billing curve. The machines cost the same asleep or at full tilt, so the marginal cost of one more overnight run is the electricity. This is also where fine-tuning lives. Training a domain adapter on your own corpus, on your own hardware, means the training data never leaves the building and the resulting weights are unambiguously yours. ## Offline, Field, and SCIF Operations The last category is the simplest: places where the cloud is not slow or expensive but absent. An AI stack that assumes connectivity is a stack that fails exactly when the environment gets interesting. - Representative workloads: in-flight development, remote field research, defense work inside a SCIF where radios are surrendered at the door, and disaster response where the network went down with everything else. - The hardware: a standalone M5 Pro or M5 Max MacBook Pro for the person, and an M6 Mac mini for the site. Both run the same models and the same tooling as the rack back home, so nothing about the workflow changes when the connectivity does. - The cloud contrast: complete operational independence. There is no degraded mode, because there is no dependency to degrade. I wrote a meaningful fraction of this series on airplanes with a local model as my pair programmer. The surprise was not that it worked. It was how little I missed the alternative. ## The enterprise decision matrix: local Mac vs. cloud Here is the whole argument in one table. Neither column wins every row, and that is the point: the goal is to route each workload to the side of the table where it belongs. Dimension Local Apple Silicon: M5 Ultra or M6 Swarm Cloud AI APIs: OpenAI, Anthropic, AWS Data sovereignty 100% on-device with zero retention by design External transmission, bound by vendor policy Cost scaling Fixed CAPEX with zero marginal token cost Linear OPEX that scales per token, every month Latency First token served off a local memory bus Network round trip plus server queue, variable Edge and physical AI Direct ROS 2 and hardware bus integration High-latency remote control loop, unsafe for actuators Max model capacity Bound by unified memory, up to 512GB on the M5 Ultra Multi-trillion parameter frontier models Execution limits Uncapped 24/7 batch execution RPM and TPM rate limits, plus vendor outages Read the capacity row carefully, because it is the cloud's strongest and most honest claim. A hosted frontier model will out-reason anything that fits in 512GB. The mistake is paying frontier prices for the 80 to 90 percent of daily traffic that never needed frontier reasoning. ## The hybrid deployment framework: local first, cloud fallback The mature architecture is not local versus cloud. It is local first, cloud fallback, with a router in front deciding which side each request deserves. In practice that means 80 to 90 percent of traffic, the routine developer queries, the agentic loops, the private RAG lookups, lands on Mac hardware you own at zero marginal cost. The cloud handles the two things it is genuinely better at: context windows beyond what unified memory holds, and the multi-trillion parameter reasoning tasks where the frontier model earns its price. The pattern holds far above the scale this series is written for. Netflix's AI platform team documented their in-house LLM serving stack, which runs open-weight and custom models on vLLM inside the same scoring service that handles their other production models, with real-time and batch paths side by side. When a company with that much negotiating leverage over API vendors still decides its LLM traffic belongs on infrastructure it operates itself, serving models on hardware you control stops being a small-team economy measure and starts looking like the default. - Route by sensitivity first. Anything touching regulated or proprietary data stays local, no exceptions. The router enforces the policy so individual engineers never have to remember it. - Route by capability second. Requests that exceed the local model's context or reasoning budget escalate to the cloud, after the sensitivity gate has already stripped what must not leave. - Let the fallback earn its keep. The cloud tier also covers hardware failures and demand spikes, which is a better job description than serving your autocomplete. The setup guide shows the cloud fallback wiring on a real cluster, and the agent hosting article covers running the always-on local tier securely. ## Frequently asked questions ### What are the best use cases for local AI on Apple silicon? Local Apple silicon wins wherever the workload is privacy-bound, latency-bound, volume-bound, or connectivity-bound. The strongest fits are air-gapped compliance work such as legal contract auditing and healthcare PHI processing, high-volume agentic coding loops that would otherwise hit API rate limits and token bills, real-time robotics and physical AI where a cloud round trip is a safety hazard, overnight batch processing of large document sets, and offline environments from airplanes to SCIFs. ### When is cloud AI still the right choice over local hardware? Cloud APIs remain the right tool when a task genuinely needs a multi-trillion parameter frontier model, a context window beyond what unified memory holds, or burst capacity far above what owned hardware provides. The practical enterprise pattern is local-first with cloud fallback: route 80 to 90 percent of routine queries, agentic loops, and private RAG to local machines, and reserve the cloud for the hardest reasoning tasks. ### What hardware do I need for air-gapped enterprise AI? The M5 Ultra Mac Studio with 512GB of unified memory at 1.2TB/s of bandwidth is the strongest single-box option. It holds a frontier-scale open-weight model plus a full enterprise RAG index in one memory pool, with no network interface required once models are loaded. Teams scaling horizontally instead deploy swarms of M6 Mac minis connected over Thunderbolt 5. ### Can Apple silicon run robotics workloads like ROS 2? Yes. ROS 2 runs on macOS, and an onboard M6 Mac mini has the memory bandwidth and the Dual 16-core Neural Engine to handle sensor fusion from LiDAR and spatial cameras while a vision-language-action model such as OpenVLA runs on the GPU. Keeping the control loop on the robot eliminates the cloud round-trip latency that makes remote inference unsafe for actuator control. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 Currently Reading ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Current Page Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Bring Zach to speak at your next event I keynote engineering conferences and executive retreats on zero-cloud AI, local inference economics, and physical AI architecture, grounded in the numbers from running Apple silicon in production. See AI Speaking Programs Bring the AI Workshop In-House --- ## Page: Local AI security: ISO 27001, SOC 2 & GDPR compliance **URL:** /projects/ai-mac-cluster/compliance **Description:** How local model inference on an Apple silicon cluster satisfies ISO 27001:2022 and SOC 2 audits, and simplifies GDPR by keeping data in-house. # Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance How we built an on-premise Apple silicon M4 cluster that satisfies rigorous enterprise data audit requirements. ## The compliance challenge in corporate AI For modern enterprise organizations, data security is the single largest blocker to artificial intelligence adoption. Transferring customer information, database structures, and private corporate intellectual property to external LLM providers exposes businesses to severe regulatory compliance penalties. Running model inference on local hardware resolves this boundary challenge completely. ## What are ISO 27001:2022 and SOC 2? ISO 27001:2022 and SOC 2 are the gold standards of information security auditing. While they share the core objective of protecting sensitive data, they approach security posture through different regulatory structures: - ISO 27001:2022: An international standard published by the International Organization for Standardization. It outlines the requirements for establishing, implementing, maintaining, and continually improving an Information Security Management System. Vetted guidelines and requirements can be found on the ISO/IEC 27001 reference page. - SOC 2 (System and Organization Controls 2): A framework developed by the American Institute of Certified Public Accountants. It audits a service organization's controls relevant to security, availability, processing integrity, confidentiality, and privacy based on Trust Services Criteria. Detailed audit criteria are available on the AICPA SOC reference page. Speaker Note: I had the privilege of delivering a keynote address in London at AICPA & CIMA Engage 2024. The professionals at that organization are truly fantastic, and their dedication to establishing clear, actionable governance and trust controls is reflected throughout their auditing standards. ## Why do companies pursue ISO 27001:2022 and SOC 2? For technology providers and enterprise partners, obtaining independent security certifications is not an academic exercise. It provides direct, high-value business benefits: - Passing Enterprise Vendor Security Reviews: Modern procurement teams mandate independent security certifications. Lacking a SOC 2 report or ISO 27001:2022 certificate immediately disqualifies tech vendors during the initial vetting phase. - Establishing Customer Trust: Audits prove that a company treats user data with maximum security. Vetted audits give corporate clients the confidence required to integrate modern AI workflows. - Institutionalizing Secure Engineering Practices: Going through a compliance audit replaces informal setups with reliable, repeatable, and automated configuration baselines across all development teams. - Mitigating Data Breach Risks: Systematically implementing audit controls hardens infrastructure against attacks, minimizing operational vulnerabilities, financial liabilities, and reputational damage. ## Where GDPR fits: local inference as a data residency strategy ISO 27001:2022 and SOC 2 are voluntary certifications. The General Data Protection Regulation is law. Any organization processing the personal data of people in the European Union is subject to it regardless of where the company is headquartered, and violations carry fines of up to 20 million euros or 4 percent of global annual revenue, whichever is higher. For AI workloads, GDPR changes the architecture conversation in a way the voluntary frameworks do not, because it regulates where personal data flows, not just how well it is protected. Sending customer records, support transcripts, or call audio to a third-party LLM API makes that provider a data processor under Article 28, which means negotiating a data processing agreement, auditing the provider's retention and logging behavior, and documenting the flow in your records of processing activities. If the provider's servers sit outside the EU, Chapter V's cross-border transfer rules stack on top: standard contractual clauses, transfer impact assessments, and an ongoing dependency on adequacy decisions that can change with a court ruling. Local inference removes that entire branch of the compliance tree. When the model runs on hardware you own, no new processor enters the picture and no international transfer occurs for that workload. The obligations that remain are ones the cluster architecture already serves: - Data Protection by Design, Article 25: Choosing an architecture where personal data never leaves the network boundary is the textbook example of privacy by design, and it is a one-sentence answer in a regulator's questionnaire. - Security of Processing, Article 32: The same controls that satisfy ISO 27001:2022 auditors, FileVault encryption at rest, restricted system users, and the hardened physical perimeter below, serve directly as Article 32 evidence of appropriate technical measures. - Right to Erasure, Article 17: Deletion requests are tractable when data lives in systems you control. There is no vendor log retention policy to chase and no third-party backup schedule to audit. - Data Minimization, Article 5: A local pipeline can transcribe, summarize, and discard raw audio in one pass on one machine, so the personal data that persists is only what the workflow actually needs. One caution: an ISO 27001:2022 certificate supports GDPR compliance but does not equal it. GDPR also governs lawful basis, consent, and data subject rights, which are legal and process questions no hardware architecture answers on its own. What the local cluster does is shrink the technical surface a data protection officer has to defend, from a chain of vendors to a locked room. ## Unique compliance challenges of on-premise Mac clusters While local AI clusters keep data entirely within physical control, deploying on-premise hardware creates unique compliance risks that are absent in public cloud platforms: - Physical Security and Asset Theft: Public clouds secure servers behind biometric entry gates and armed personnel. An on-premise hardware setup is vulnerable to physical tampering, unauthorized local device access, or direct system theft. - Configuration Drift across Nodes: Without automated cloud hypervisors, managing separate physical machines risks manual configuration variance. System updates, security patches, and OS-level configurations must be kept identical to satisfy audits. - Lack of Centralized Audit Logging: Unlike public clouds with built-in telemetry, physical nodes generate separate local system logs. Standard compliance controls require proving that unauthorized login attempts or administrative tasks are captured and centralized. - Logical Access Control and Lateral Movement: Running diverse agentic workloads on local systems raises isolation risks. Without strict user boundary enforcement, a compromised execution script could gain root privileges and access files across the entire cluster. ## Mitigating network and access risks: the hardened physical data perimeter To mitigate localized hardware vulnerabilities, physical security risks, and logical boundary gaps within the cluster, we implement a comprehensive physical data perimeter: - Activating the System Firewall: We enable the built-in macOS application firewall on every node, blocking unauthorized inbound connections and establishing immediate network-level defense boundaries. - Deactivating Background Services and Protocols: To minimize the network attack surface, we completely disable unused pre-installed background services such as AirDrop, Wi-Fi, and Bluetooth, ensuring all node communications are routed strictly over wired network interfaces. - Uninstalling Unnecessary Applications: We purge pre-installed apps and non-essential system software, leaving only a minimalist, highly secure footprint dedicated entirely to running model inference tasks. - Managing Physical Security and Telemetry: Deploying localized hardware shifts physical security directly to our administration. In addition to securing physical machine access, we monitor comprehensive physical environment metrics including node fan speeds and core operating temperatures rather than just traditional CPU and RAM usage. - Executing Services Under Unprivileged Users: Every running application executes under restricted, non-administrator user accounts configured to perform only their dedicated application processes. This architecture guarantees that even if a service is compromised, the attacker lacks the system-level permissions required to modify OS configurations. - Centralizing Telemetry Logs in the Cloud: System audit logs and runtime events are streamed continuously to a secure cloud platform. If a physical node or the entire local cluster experiences a catastrophic power outage or goes offline, we retain the complete operational history needed to perform forensic security investigations. ## Mitigating theft and tampering: disk encryption and ephemeral processing To mitigate physical security risks and the threat of physical theft, we enforce full disk encryption combined with a strictly stateless compute model: - FileVault Disk Encryption & Lockdown: Physical security is reinforced by enforcing full FileVault disk encryption on every Mac mini. Any physical disconnection or power outage forces an immediate shutdown, locking the encrypted volumes and preventing data recovery without the administrative security keys. - Ephemeral Processing Model: The Mac minis in our cluster serve strictly as stateless compute workers. They pull transcription or inference payloads from an AWS SQS queue, process the workload in active RAM, and immediately push the resulting output to a secure cloud API hosted on AWS over an encrypted HTTPS connection. Once the task completes, the local system completely clears the temporary workspace, leaving zero persistent customer data on local drives. ## Mitigating configuration drift: satisfying ISO 27001:2022 and SOC 2 audits To eliminate configuration drift and telemetry gaps, all nodes are configured from a hardened base image and audited continuously: - Continuous Vulnerability Monitoring: The physical macOS nodes are treated strictly as production servers rather than general office workstations. They are scanned continuously using Tenable agents and Intruder configuration audits to identify and patch system level issues promptly. - Standardized Base Image Hardening: To maintain a clean security posture, unnecessary background services are disabled, unused pre-installed applications are removed, and all nodes are provisioned starting from a hardened base system image snapshot. During our recent corporate security audits, we successfully proved that localized AI models completely eliminate data-in-transit compliance vulnerabilities. Because the physical hardware resides inside our audited perimeter, acts purely as an ephemeral processing layer, and secures active drives under FileVault encryption, we demonstrated complete control over customer data, meeting all necessary SOC 2 and ISO 27001:2022 audit controls without exception. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 Currently Reading ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Current Page Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Need help with local AI security compliance? I speak to engineering and compliance leaders about secure local AI architecture, and run workshops that walk teams through the audit posture in this guide, from network boundaries to ISO 27001:2022, SOC 2, and GDPR readiness. See AI Speaking Programs Run the 60-Minute Security Audit --- ## Page: Apple silicon LLM memory & speed calculator: M1 to M6 **URL:** /projects/ai-mac-cluster/llm-speed-calculator **Description:** Free calculator for local LLMs on Apple silicon. Pick a chip from M1 to M6, model size, context, and quantization to see RAM and tokens per second. # Apple silicon local LLM memory and speed calculator Pick a chip from M1 to M6, a model size, a context length, and a quantization level. Get the unified memory you need, the Mac configuration that fits, and theoretical prefill and generation speeds. ## Size your local LLM setup Every Apple silicon Mac shares one pool of unified memory between the CPU and GPU, so the question of whether a model runs locally comes down to arithmetic: model weights plus KV cache plus runtime overhead, measured against the slice of RAM macOS lets the GPU claim. Generation speed follows memory bandwidth almost linearly. This calculator does that arithmetic for every chip Apple has shipped, from the original M1 to the M5 Ultra and M6. Apple silicon chip Model size Context length Quantization Results update as you change a selection. Assumptions: FP16 KV cache, 70 percent memory bandwidth efficiency for generation, 25 percent compute efficiency for prefill, and 2 GB of runtime overhead. ### Memory Requirements Model weights 42.8 GB KV cache at 16k context 5.4 GB Runtime overhead 2.0 GB Required unified memory 50.2 GB Fits on the M4 Max with 128 GB of unified memory. About 96 GB of that is available to the GPU under the default macOS limit. ### Theoretical Speed Generation speed with a short prompt ~8.9 tokens/sec Generation speed at the full 16k window ~7.9 tokens/sec Prompt prefill speed ~65.2 tokens/sec Time to first token at full context ~4.2 minutes Speeds are theoretical estimates for llama.cpp or MLX class runtimes. M5 family and M6 prefill figures assume the GPU neural accelerators and are directional until independent benchmarks land. ## How the math works ### Required unified memory The memory a local model needs is the sum of three parts. The weights take parameter count times bits per weight divided by eight, so a 70B model at Q4_K_M is 70.6 billion times 4.85 bits, or about 42.8 GB. The KV cache grows linearly with context: two tensors per layer, times the number of KV heads, times the head dimension, times two bytes per FP16 value, times your context length. Runtime overhead for the inference engine, compute buffers, and the operating system adds roughly 2 GB more. macOS does not hand the GPU the whole memory pool. The default working set limit sits near 75 percent of unified memory on higher-memory machines and closer to two thirds on smaller ones, which is why a 47 GB model does not fit on a 48 GB Mac. You can raise the ceiling with a sysctl, at the cost of squeezing the rest of the system: sudo sysctl iogpu.wired_limit_mb=57344 ### Generation speed is a bandwidth problem Generating one token requires reading every active weight from memory once, so the theoretical ceiling is memory bandwidth divided by the size of the active weights. A 70B model at Q4_K_M reads 42.8 GB per token: at the M4 Max's 546 GB/s that caps out near 12.8 tokens per second, and real runtimes deliver about 70 percent of the ceiling. This is why the M5 Ultra's 1.2 TB/s matters more for local LLMs than any GPU core count, and why the calculator also charges the growing KV cache against bandwidth as your context fills up. ### Prefill speed is a compute problem Before the first token appears, the model processes your entire prompt. That phase batches many tokens through the weights at once, so it is limited by raw matrix multiply throughput rather than bandwidth: roughly two floating point operations per active parameter per prompt token. The calculator assumes 25 percent of peak FP16 throughput, which matches what Metal backends achieve in practice. The M5 family changes this equation materially, because its GPU neural accelerators multiply matrices far faster than the M4 generation, cutting the long wait before the first token on big prompts. ## Apple silicon memory bandwidth, M1 through M6 Bandwidth figures and memory ceilings are Apple's published specifications. The FP16 column is an estimate from public GPU core counts, and the last column shows the largest model in this calculator that fits the top memory configuration at Q4_K_M with an 8k context. The M5 Ultra pairs 1.2 TB/s with up to 512 GB of unified memory, while the M6 launches as a base chip at 170 GB/s with no Pro, Max, or Ultra variants announced. Chip Memory Bandwidth Max Unified Memory FP16 Compute Largest Model at Q4_K_M M1 68 GB/s 16 GB 5.2 TFLOPS 8B M1 Pro 200 GB/s 32 GB 10.4 TFLOPS 8B M1 Max 400 GB/s 64 GB 21 TFLOPS 70B M1 Ultra 800 GB/s 128 GB 42 TFLOPS 70B M2 100 GB/s 24 GB 7.2 TFLOPS 8B M2 Pro 200 GB/s 32 GB 13.6 TFLOPS 8B M2 Max 400 GB/s 96 GB 27.2 TFLOPS 70B M2 Ultra 800 GB/s 192 GB 54.4 TFLOPS 70B M3 100 GB/s 24 GB 8.2 TFLOPS 8B M3 Pro 150 GB/s 36 GB 12.8 TFLOPS 32B M3 Max 400 GB/s 128 GB 28.4 TFLOPS 70B M3 Ultra 819 GB/s 512 GB 56.8 TFLOPS 235B MoE M4 120 GB/s 32 GB 9.2 TFLOPS 8B M4 Pro 273 GB/s 64 GB 18.4 TFLOPS 70B M4 Max 546 GB/s 128 GB 36.8 TFLOPS 70B M5 153 GB/s 32 GB 37 TFLOPS est. 8B M5 Pro 307 GB/s 64 GB 74 TFLOPS est. 70B M5 Max 614 GB/s 128 GB 148 TFLOPS est. 70B M5 Ultra 1228 GB/s 512 GB 296 TFLOPS est. 235B MoE M6 170 GB/s 32 GB 45 TFLOPS est. 8B ## How much RAM each model size needs The short answer for the most common question: a 70B LLM on a Mac needs a 64 GB machine at minimum, and a 128 GB machine to be comfortable. The table below sizes each model class at Q4_K_M with an 8k context, against the default macOS GPU memory limit. Larger context windows shift these numbers up, which is exactly what the calculator above is for. Model Q4_K_M Weights Required at 8k Context Minimum Mac RAM 8B dense, Llama 3.1 8B or Qwen3 8B class 4.8 GB 7.9 GB 16 GB 32B dense, Qwen3 32B class 19.9 GB 24.0 GB 36 GB 70B dense, Llama 3.3 70B class 42.8 GB 47.5 GB 64 GB 235B MoE, Qwen3-235B-A22B 142.5 GB 146.0 GB 256 GB Model architecture details come from the published GGUF releases of each reference model, and the sizes match what llama.cpp reports when it loads them. Bandwidth figures for the newest chips come from Apple's M6 and M5 Ultra announcement. ## Frequently asked questions ### How much RAM do I need to run a 70B LLM on a Mac? A 70B dense model at Q4_K_M quantization needs about 42.8 GB for the weights alone, and roughly 47.5 GB once you add an 8k KV cache and runtime overhead. The smallest Apple silicon configuration that hosts it under the default macOS GPU memory limit is 64 GB of unified memory, and it is a tight fit. For 32k context windows or Q8_0 quality, plan on 128 GB. ### How many tokens per second does the M5 Ultra generate? The M5 Ultra moves 1.2 TB/s of unified memory bandwidth, so at a 70 percent efficiency assumption it generates about 20 tokens per second on a 70B dense model at Q4_K_M, and about 64 tokens per second on Qwen3-235B-A22B, because the MoE model only reads 22B active parameters per token. ### Can a Mac mini run a 70B model? Yes, narrowly. A Mac mini with an M4 Pro and 64 GB of unified memory fits a 70B model at Q4_K_M with a short context window, but at 273 GB/s of bandwidth it generates only about 4.5 tokens per second. That is usable for batch jobs and unattended agents, not for interactive chat. For a responsive 70B experience, a Mac Studio with a Max or Ultra class chip is the better host. ### What is the difference between Q4_K_M and Q8_0? Q4_K_M stores weights at about 4.85 bits each and Q8_0 at about 8.5 bits, so Q8_0 roughly doubles the memory footprint and halves the generation speed on the same hardware. Q8_0 is nearly indistinguishable from the FP16 original, while Q4_K_M gives up a small amount of quality that most workloads never notice. Start at Q4_K_M and move up only if you measure a quality problem. ### Why does a 235B MoE model run faster than a 70B dense model? A mixture of experts model stores every expert in memory but activates only a few per token. Qwen3-235B-A22B holds 235B parameters in RAM yet reads just 22B per generated token, so it needs the memory of a giant model but generates at the speed of a 22B model. Memory capacity requirements follow total parameters, and speed follows active parameters. ## Put the numbers to work This calculator exists because I run local models in production. The Mac cluster series covers what happens after you pick the hardware: building an M4 Mac mini cluster that cut our cloud AI spend by $40k per year, the step-by-step cluster setup guide, running local agentic coding on Apple silicon, and hosting always-on private AI agents on a Mac mini. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. ### Planning a Local AI Deployment? I keynote conferences on the hardware math behind local AI, and my hands-on workshop takes engineering teams from sizing the machines to shipping a private agentic stack. See AI Speaking Programs Bring the AI Workshop In-House --- ## Page: Lightwall AI interactive installation | Zach Rattner **URL:** /projects/lightwall **Description:** Experience Lightwall: a fusion of kinetic light, fine art, and real-time spatial AI that actively responds to human presence and voice. # Lightwall ## Experience art that reacts to you Watch Demo The goal was simple but ambitious: to teach a machine to feel. Lightwall is a collaboration between Zach Rattner and artist Rita Sus that merges fine art, kinetic light, and real-time AI to create a living, breathing environment shaped by the people who stand before it. It bridges the gap between digital intelligence and physical reality, turning a static wall into a living presence. Unlike standard static installations, this composition responds to the visitor's movement and voice. When approached, the artwork shifts from an abstract array of geometric shapes and glass cubes to an entity actively engaging in dialogue. Lightwall premiered at the California Center for the Arts Museum and is now being prepared to exhibit at Dutch Design Week in Eindhoven in October 2026. ### Collaborators #### Lead Artist Rita Sus #### Technical Director Zach Rattner ### Technologies - LEDs - Motors - Radar - Microphone - Embedded Computer ## Origin & Process Rita and I sought to merge the tangible beauty of kinetic glass with the invisible power of real-time AI. The piece evolved from experimental prototypes into something that feels alive rather than assembled. A sign placed next to the piece simply invites visitors: "Talk to the wall. It's listening." When approached, Lightwall speaks back in a robotic, yet distinctly engaging voice. As recently featured in the Orange County Register, the installation actively invites users into an esoteric conversation about its purpose and identity. ## The interaction model Most AI experiences are fundamentally passive. You type a prompt into an empty text box on your phone and wait for the screen to blink back a response. We wanted to physically invert that relationship. Instead of the human looking inside the machine, we designed the machine to look out at the human. Lightwall uses an advanced array of visual and acoustic sensors to map its immediate physical environment in real time. ## Philosophy: private and offline Lightwall is designed to see you, not to watch you. One of our design constraints was treating the user's data respectfully. By running entirely on-site off an inference cluster with no cloud connection, the installation creates a fleeting, private dialogue. Just like with a person. - No traces left behind: It perceives movement, reacts instantly, and forgets. Every conversation dies the moment the visitor walks away. - Highly portable: Because of the standalone inference architecture, the piece is entirely self-contained. It requires only standard AC power. Lightwall is about using AI not to replace creativity, but to make human experiences richer. If you're interested in booking Lightwall for an installation or learning more about the intersection of kinetic art and embedded AI, reach out to the team. Watch Demo --- ## Page: Scourhead: AI web research agent | Zach Rattner **URL:** /projects/scourhead **Description:** Meet Scourhead: a free, open-source AI agent designed to streamline online research by scouring the web and organizing data into spreadsheets. # Scourhead ## AI research agent A free, open-source AI agent that streamlines your online research. Scourhead autonomously scours the web, organizes findings, and delivers the data directly into a spreadsheet—running entirely locally on Windows, macOS, or Linux. Book Zach to speak on AI Agents View GitHub Researching the web for specific data points can be incredibly tedious. Whether you're building a lead list, finding competitor pricing, or gathering contact information, the process usually involves hours of clicking, copying, and pasting into a spreadsheet. Scourhead is an open-source AI agent designed to automate this entirely. By taking a simple natural language prompt and a spreadsheet template, Scourhead navigates the internet on your behalf. It visits websites, extracts the relevant information, and structures it perfectly into your spreadsheet. Best of all, Scourhead is designed to run locally on your own hardware without relying on expensive cloud subscriptions. It's available for Windows, macOS, and Linux, ensuring your data remains private and your operational costs stay at zero. ### Features - Autonomous web searching - Automated data extraction - Direct spreadsheet integration - 100% free & open-source - Cross-platform (Mac, Windows, Linux) ### Pioneering agentic workflows I built and open-sourced Scourhead in 2024, well before "AI Agents" became a corporate buzzword. It served as a vital proof-of-concept for how AI could move beyond generating text to actually taking action. Today, the enterprise market is flooded with AI agents, but Scourhead remains a foundational example of how localized, privacy-first AI automation transforms the future of work. ## Privacy & Security Scourhead runs locally on your machine. There are no mandatory cloud accounts or subscriptions. You control the data, the execution environment, and the final output. ## Why this matters for business leaders Scourhead offers a glimpse into the future of the workforce. It proves that tasks previously requiring dozens of human hours—like lead generation and competitor research—can now be fully automated at zero operational cost. In my keynotes, I use the lessons learned from building Scourhead to teach executives how to identify "agent-ready" workflows within their own organizations. ### Ready to build your own agents? Want to teach your engineering team how to build tools like Scourhead? Book my Modernize Your Engineering Org for the AI Era workshop for a hands-on session on AI automation. View Workshop Details ## Community & Press Scourhead sparked a massive conversation in the automation community, reaching tens of thousands of professionals across various networks. Though an experimental release, it was heavily featured across numerous tech directories, blogs, and social feeds: - Aiaxio - Best of AI - Nabin Chalise (Facebook) - Peachey Publications (LinkedIn) - Slashdot - Softonic - SourceForge - Tenere Team - Toolify.ai --- ## Page: 31 granted US patents | Zach Rattner **URL:** /patents **Description:** Zach Rattner's 31 granted US patents across artificial intelligence, computer vision, consumer electronics, and telecommunications. # Top AI keynote speaker: a history of innovation I build artificial intelligence for a living. The stories on stage come from shipping it. As the CTO and Co-Founder of Yembo, I've deployed AI around the world, helping businesses complete more inspections every day. I bring the battle scars of a technical founder directly to the stage. Technology only matters when it makes a difference in the hands of those using it. 1/6/2026 ## System and Method for an Enhanced Hair Dryer US 12,514,353 ### The Invention Smart hardware that utilizes environmental sensors to reduce energy consumption and customize output. ### The Takeaway Transforms passive analog tools into active data-gathering assets by leveraging integrated environmental telemetry algorithms. Inventors: Ryan Goldman , Zach Rattner , Jonathan Friedman 12/23/2025 ## System and Method for an Enhanced Hair Dryer US 12,501,981 ### The Invention Integration of cloud-based storage with proximal device sensors for rapid data processing. ### The Takeaway Accelerates local data processing by bridging cloud-based storage directly with proximal environmental device sensors. Inventors: Ryan Goldman , Zach Rattner , Jonathan Friedman 9/30/2025 ## System and Method for an Enhanced Hair Dryer US 12,426,696 ### The Invention Refined connectivity functions within consumer hardware to ensure seamless wireless local area network transmission. ### The Takeaway Refines internal connectivity logic to ensure constant, seamless wireless network transmission in edge-deployed hardware. Inventors: Ryan Goldman , Zach Rattner , Jonathan Friedman 8/26/2025 ## Anonymizing Personally Identifying Information in Image Data US 12,400,033 ### The Invention System that automatically scrubs personally identifying information (PII) from visual data to solve a massive hurdle for enterprise AI. ### The Takeaway Resolves privacy conflicts by automatically detecting and permanently anonymizing personally identifying information before visual data is processed. Inventors: Zach Rattner , Devin Waltman , Noel Kennebeck , Maciej Halber 8/19/2025 ## System and Method for an Enhanced Hair Dryer US 12,389,996 ### The Invention Optimized hardware energy usage to enable highly efficient product designs. ### The Takeaway Reduces overall hardware energy usage through algorithmic sensor processing, enabling highly efficient, sustainable product designs. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 8/5/2025 ## System and Method for an Enhanced Hair Dryer US 12,376,661 ### The Invention Integration of infrared and temperature camera sensors directly into consumer-grade hardware. ### The Takeaway Embeds complex infrared and temperature camera sensors directly into consumer goods for live environmental mapping. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 12/31/2024 ## System and Method for an Enhanced Hair Dryer US 12,178,303 ### The Invention Core logic for determining user-specific, customizable settings based on real-time environmental reads. ### The Takeaway Analyzes real-time environmental reads to automatically deploy personalized, highly accurate hardware adjustments for end users. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 12/10/2024 ## Systems and Methods for Delivering Head in a Battery-Powered Hair Dryer US 12,161,209 ### The Invention Novel heating element and battery cell configuration for high-output hardware. ### The Takeaway Introduces a novel heating element design that successfully pairs high-drain thermal requirements with standard portable battery packs. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 5/21/2024 ## Browser Optimized Interactive Electronic Model Based Determination of Attributes of a Structure US 11,989,848 ### The Invention System enabling heavy, interactive 3D spatial models to run natively within web browsers, circumventing the need for expensive native applications. ### The Takeaway Optimizes processing loads so anyone can interactively verify and modify complex 3D environments natively within web browsers. Inventors: Kyle Babinowich , Maciej Halber , Marc Eder , Janpreet Singh , Zach Rattner 4/2/2024 ## Modulation Techniques for Prolonging Battery Life in a Battery-Powered Hair Dryer US 11,944,176 ### The Invention Modulation techniques that extend battery runtime without degrading performance. ### The Takeaway Implements intelligent modulation techniques for infrared heating elements, substantially extending battery runtime and optimizing output across the lifecycle. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 9/12/2023 ## System and Method for an Enhanced Hair Dryer US 11,751,657 ### The Invention System for storing and recalling specific user profiles to deliver coordinated, optimized hardware solutions. ### The Takeaway Stores and recalls individualized user environmental profiles to instantly customize and optimize hardware performance on demand. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 5/23/2023 ## Systems and Methods for Building a Virtual Representation of a Location US 11,657,419 ### The Invention Machine learning and geometric models required to annotate 3D models with semantic spatial information in real-time. ### The Takeaway Automatically processes unstructured 2D images into fully annotated 3D environments, permanently linking semantic metadata to physical objects. Inventors: Marc Eder , Siddharth Mohan , Maciej Halber , Anoop Jakka , Devin Waltman , Zach Rattner 5/23/2023 ## Capacity Optimized Electronic Model Based Prediction of Changing Physical Hazards and Inventory Items US 11,657,418 ### The Invention Automated extraction of hazard data and layout structure to materially improve insurance underwriting accuracy. ### The Takeaway Automates the ingestion of physical hazard data, enabling accurate insurance underwriting without deploying human inspectors on-site. Inventors: Devin Waltman , Siddharth Mohan , Anoop Jakka , Marc Eder , Maciej Halber , Zach Rattner 2/21/2023 ## Systems and Methods for Cooling Batteries in a Battery Powered Blow Dryer US 11,583,052 ### The Invention Dual-purpose cooling feature to maintain hardware integrity during high-performance output. ### The Takeaway Engineers a dual-purpose cooling loop that simultaneously lowers battery temperatures and augments primary physical hardware performance. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 12/6/2022 ## Identifying Flood Damage to an Indoor Environment Using a Virtual Representation US 11,521,273 ### The Invention Computer vision system to automatically assess and categorize environmental damage based on virtual reference planes. ### The Takeaway Uses virtual reference planes across 3D models to instantly calculate damage extent, making claims processing faster and objective. Inventors: Anoop Jakka , Siddharth Mohan , Devin Waltman , Marc Eder , Maciej Halber , Zach Rattner 5/17/2022 ## Artificial Intelligence Generation of an Itemized Property and Renters Insurance Inventory List for Communication to a Property and Renters Insurance Company US 11,334,901 ### The Invention Computer vision architecture that powers Yembo, automating the generation of itemized statements of work based purely on sensor data. ### The Takeaway Powers comprehensive computer vision architectures that instantly translate multi-sensor room data into actionable commercial quotes. Inventors: Zach Rattner , Siddharth Mohan 3/8/2022 ## Systems and Methods for Providing AI-Based Cost Estimates for Services US 11,270,363 ### The Invention Machine learning system to instantly generate interactive quotes for physical services based on visual context data. ### The Takeaway Utilizes machine learning models to instantly convert complex physical location descriptions into interactive, dynamic service quotes. Inventors: Zach Rattner , Siddharth Mohan , Vikram Gupta 3/1/2022 ## Schema Translation Systems and Methods US 11,263,460 ### The Invention Infrastructure for real-time video feed analysis, where machine learning models identify and index objects live. ### The Takeaway Enables real-time video feed analysis, immediately identifying, indexing, and categorizing objects the moment they appear on camera. Inventors: Zach Rattner , Tiana Hayden , Noel Kennebeck 6/15/2021 ## Systems and Methods for Cooling Batteries in a Battery Powered Blow Dryer US 11,033,089 ### The Invention Advanced thermal management systems for portable hardware. ### The Takeaway Advances thermal management systems to protect highly stressed portable hardware components without sacrificing operational output. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 5/11/2021 ## Systems and Methods for Delivering Heat in a Battery Powered Blow Dryer US 11,000,104 ### The Invention Infrared technology to generate high output with reduced power consumption. ### The Takeaway Replaces traditional resistive heating with infrared technology to generate high thermal output with substantially reduced power draw. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 12/15/2020 ## Systems and Methods for Providing AI-Based Cost Estimates for Services US 10,867,328 ### The Invention Foundational logic for linking visual data directly to financial estimation models. ### The Takeaway Bridges visual context and financial forecasting by automatically generating itemized service quotes directly from raw visual data. Inventors: Zach Rattner , Siddharth Mohan , Vikram Gupta 6/23/2020 ## System and Method for an Enhanced Hair Dryer US 10,687,597 ### The Invention Core proprietary calculations (V-Factor) for syncing hardware performance with consumable products. ### The Takeaway Utilizes proprietary V-Factor calculations to synchronize connected hardware performance directly with consumable product profiles. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 6/23/2020 ## Systems and Methods for Delivering Heat in a Battery Powered Blow Dryer US 10,687,596 ### The Invention Unique battery cell configuration required to maintain high heat output in untethered devices. ### The Takeaway Employs a unique battery cell configuration that safely maintains high heat output in untethered portable environments. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 3/31/2020 ## Systems and Methods for Cooling Batteries in a Battery Powered Blow Dryer US 10,608,297 ### The Invention Baseline patents for the thermodynamics of portable, high-draw consumer devices. ### The Takeaway Captures and repurposes cooling exhaust to maintain battery integrity while actively improving the device's primary function. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 1/7/2020 ## Artificial Intelligence Prediction Algorithm for Generating an Itemized Statement of Work and Quote for Home Services Based on Two-Dimensional Images, Text, and Audio US 10,528,962 ### The Invention Ingestion of multimodal data (2D images, text, and audio) into a single predictive algorithm for the home services sector. ### The Takeaway Unifies multi-sensor spatial data, allowing enterprises to automatically generate accurate, itemized statements of work from environmental scans. Inventors: Zach Rattner , Siddharth Mohan 9/10/2019 ## Systems and Methods for Delivering Heat in a Battery Powered Blow Dryer US 10,405,630 ### The Invention Baseline methodology for infrared integration in portable consumer goods. ### The Takeaway Overcomes portable power limitations, enabling high-heat battery output via efficient infrared elements and optimized cell arrangements. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 7/23/2019 ## Apparatuses and Methods for Fast Onboarding an Internet-Enabled Device US 10,360,362 ### The Invention Reduces friction for users adopting new connected IoT devices. ### The Takeaway Automates secure onboarding by allowing administrative devices to seamlessly authenticate and pass network credentials to headless nodes. Inventors: Zach Rattner , Christopher Wingert , Rob Daley , Patrik Lundqvist 7/17/2018 ## System and Method for an Enhanced Hair Dryer US 10,021,952 ### The Invention Foundational IP for a smart, connected apparatus that bridges the gap between digital data and physical styling. ### The Takeaway Integrates IoT connectivity and sensors into hardware, enabling dynamic environmental adjustments that significantly reduce power consumption. Inventors: Ryan Goldman , Jonathan Friedman , Zach Rattner 4/4/2017 ## Electromagnetic Mating Interface US 9,613,739 ### The Invention System using electromagnetic forces to automatically couple and exchange data between modular physical devices. ### The Takeaway Uses precisely-timed electromagnetic forces to automatically couple and exchange data between units, enabling adaptable, modular hardware ecosystems. Inventors: Zach Rattner , Clayton Dumstorff , Daniel Ervin , Rob Daley 10/11/2016 ## Network Access and Control for Mobile Devices US 9,467,453 ### The Invention Secure content controls implemented directly at the modem level of mobile devices to prevent unauthorized access. ### The Takeaway Moves content authorization controls to the physical modem level, ensuring strict, immutable data compliance regardless of software state. Inventors: Michael Canoy , Michael DeVico , Zach Rattner , Steve Sprigg 9/13/2016 ## Radio-Agnostic Message Translation Service US 9,445,248 ### The Invention Foundation for seamless communication and translation across different, previously incompatible network technologies. ### The Takeaway Acts as a radio-agnostic translation hub, enabling disparate hardware architectures and legacy networks to communicate seamlessly. Inventor: Zach Rattner ## Bring proven AI blueprints to your next event Give your audience more than just theory. Bring a battle-tested technical founder to your next conference, and let them walk away with frameworks driving innovation in 20+ countries. Invite Zach to Speak --- ## Page: AI for construction: keynotes and workshops **URL:** /ai-for-construction **Description:** Keynotes and workshops on AI in construction and remodeling, from the CTO who turns phone video of a property into measurements and scope. # AI for construction Nearly every AEC firm already using AI plans to buy more of it. Only 27 percent use it at all. That gap is a data problem, and it starts on the job site. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Contractors & AEC ## Fees from $10,000+ Book Now ## Everyone believes. Almost nobody has shipped. Only 27 percent of AEC firms use AI for automation, problem-solving, or decision-making, while 94 percent of the firms already using it plan to increase their investment. Skepticism would show up as abandonment, and the adopters are doubling down. If your firm sits in the other 73 percent, you have plenty of company. You are the curve. The usual explanations are conservatism and thin margins, and I do not buy either one. This industry adopted laser levels, GPS grade control, and prefab quickly once each one paid for itself. Contractors adopt what works. So here is what I think is really going on. Bad data was estimated to cost global construction $1.85 trillion in a single year, including nearly $89 billion in avoidable rework. Job site reality lives in photos nobody labeled, a superintendent's memory, and a text thread with forty people in it. You cannot point a model at that and expect an answer. Your last AI pilot failed because your site was never captured in a form a machine could read. The model was fine. ## That is the problem I work on At Yembo I built ScanMyHome: walk a space with an ordinary phone and get back a measurement-grade 3D reconstruction, floor plan, dimensions, wall cutouts, and the objects in it identified. Capturing an uncontrolled space in bad light with someone's belongings in the way is the same engineering problem on a job site as in a living room, and I have spent a decade losing arguments with it and eventually winning some. Insurance carriers settle claims against those numbers. That is the capture itself rather than a dashboard sitting on top of data somebody else collected, and it is the step your industry keeps skipping and then blaming the model for. You are also short people. The industry needs something like 349,000 net new workers this year, and most firms cannot find them. That makes every avoidable site visit and every rework hour more expensive than it was two years ago, which is the business case, and notice it does not require anyone on your payroll to be replaced. What does not carry over is everything downstream of the capture: means and methods, sequencing, code, and the judgment your superintendents apply when a wall opens up and the drawing was wrong. I am not going to pretend a model helps you there. So the value I bring your team is calibration. When a vendor tells your firm what their computer vision will do next quarter, I can tell you from having built it whether that is a roadmap or a wish. If your team wants to work rather than listen, the agentic workflows workshop runs leadership teams through their own processes, and my work on AI for interior design covers the same capture problem on the design side of a remodel. ## What your audience walks out with - An honest read on why their last pilot stalled, and whether it was the tool or the data underneath it. - Which site information is worth capturing systematically, and which is fine left in a superintendent's head. - Where computer vision cuts a trip to a site, and where the trip is still the cheapest answer. - How to tell, before signing, whether a vendor's computer vision claim is a roadmap or a wish. ## Where I have taught and spoken ## Bring this to your firms Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak at construction industry events? Yes. Zach Rattner speaks to general contractors, remodelers, and architecture and engineering firms on where AI works on a real job site. His talks cover capturing site conditions from video, reducing rework, and why most construction AI pilots stall before they reach production. Why do construction AI projects fail? Zach Rattner argues that construction AI usually fails on data rather than on models. A Bluebeam survey of over 1,000 AEC professionals found only 27 percent of firms use AI for automation, problem-solving, or decision-making, while 94 percent of the firms already using it plan to increase their investment, and bad data was estimated to cost global construction $1.85 trillion in 2020 including $88.69 billion in avoidable rework. A model cannot reason about site conditions that were never captured in a form it can read. What does Zach Rattner know about capturing site conditions? Zach Rattner is the co-founder and Chief Technology Officer of Yembo, an AI platform used in more than 20 countries that turns phone video of a property into measurements, inventories, and scope without an on-site visit. He holds 31 granted US patents, many covering computer vision applied to measuring interior spaces. Can AI capture site conditions accurately enough to estimate from? Yes. ScanMyHome, which Zach Rattner built at Yembo, produces measurement-grade 3D reconstructions of interior spaces from an ordinary phone walkthrough, including floor plans, dimensions, and wall cutouts, and insurance carriers settle claims against those measurements. Zach Rattner is direct with construction audiences about where that capability stops, which is means and methods, sequencing, and code judgment. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run workshops for construction leadership teams? Yes. Alongside keynotes, Zach Rattner runs working sessions for leadership teams of 8 to 25 people, where the group maps its own estimating, site documentation, and change order workflows and leaves with a ranked list of what to automate first and what to leave alone. --- ## Page: AI for customer service and contact centers **URL:** /ai-for-customer-service **Description:** Keynotes and workshops on AI in customer service, on why deflection is the wrong metric and what resolution-first automation actually requires. # AI for customer service Deflection counts the customers you got rid of. Your customers can tell the difference, and so can your CSAT. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience CX & Support Leaders ## Fees from $10,000+ Book Now ## You are optimizing the wrong number Median tier-one deflection across enterprise CX programs sits around 41 percent in 2026. Read one layer down and the number falls apart: password resets and refunds deflect above 70 percent, while nuanced complaints rarely break 25. So a deflection rate can climb beautifully while the only thing that actually happened is that the easy contacts left and the hard ones stayed, now arriving at your agents angrier because they tried the bot first. The satisfaction data says the same thing from the other side. Pure AI handling lands around 4.1 out of 5 against 4.3 for humans, but hybrid flows with a clean escalation path close that gap to roughly five hundredths of a point. AI actually beats the human baseline when it resolves. It loses when it deflects. Deflection counts the customers you got rid of. Resolution counts the problems you solved. Only one of those is worth putting on a dashboard. ## Why the handoff is the whole game Around 74 percent of consumers find repeating themselves very frustrating, and a majority simply give up when made to do it more than once. That is the moment automation either earns its budget or costs you the customer, and it is an architecture decision rather than a model decision. This is the part I have actually built. At Yembo I have spent a decade deciding what a machine handles, what a person handles, and how the handoff between them carries context instead of dropping it. That question does not change much between a claims queue and a support queue. When a leadership team wants to work rather than listen, Turn Your Call Center Into an AI Asset takes them through build versus buy, vendor evaluation, and moving from a sampled review of calls to full coverage. You can also run your own conversations through the free Call Center Analyzer before booking anything. ## What your audience walks out with - A metric set built on resolution rather than deflection, and the argument for changing it with a CFO. - Which intents are genuinely safe to automate today, and which ones punish you for trying. - An escalation design that carries context, so nobody has to repeat themselves to a human. - A realistic 90-day view of a first deployment, and the intents to leave alone until the handoff is solid. ## Where I have taught and spoken ## Bring this to your team Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak about AI in customer service? Yes. Zach Rattner speaks to CX leaders, contact center operators, and support organizations on where conversational AI genuinely resolves customer problems and where it simply moves them out of the queue. He also runs a full workshop on the same subject for leadership teams. Does AI hurt customer satisfaction in contact centers? Industry measurement in 2026 puts pure AI handling at roughly 4.1 out of 5 CSAT against 4.3 for human agents, but hybrid flows with clean escalation narrow that gap to about 0.05 points, and AI satisfaction exceeds the human baseline when the AI resolves an issue rather than deflecting it. Zach Rattner argues that satisfaction tracks resolution rather than automation, which means the design of the escalation path matters more than the model. What is wrong with measuring contact deflection? Zach Rattner argues that deflection counts the customers a contact center got rid of rather than the problems it solved. Median tier-one deflection sits around 41 percent in 2026, but refund and password intents deflect above 70 percent while nuanced complaints rarely exceed 25 percent, so a rising deflection rate can simply mean the easy contacts moved and the hard ones stayed. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run a customer service workshop? Yes. Turn Your Call Center Into an AI Asset is a half-day or full-day workshop for groups of 8 to 25 people, covering build versus buy decisions, quality management vendor evaluation, and moving from sampled call review to full conversation coverage. --- ## Page: AI data security: keynotes and workshops **URL:** /ai-for-data-security **Description:** Keynotes and workshops on AI data security and shadow AI, from a CTO who took an AI platform through ISO 27001, SOC 2 Type II, GDPR, and NIST 800-171. # AI and data security Your AI ban made AI use invisible, and invisible is the expensive kind. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Executives & Boards ## Fees from $10,000+ Book Now ## Prohibition is what produces the breach More than half of the workers already using generative AI use tools their employer has not approved, in a Salesforce survey of more than 14,000 workers across 14 countries. Much of that usage runs through personal accounts, which is to say outside every control you have. The consequences are now measurable. IBM's Cost of a Data Breach Report found shadow AI involved in one in five data breaches, adding an average of $670,000 to the cost of each one, and 63 percent of the organizations it hit had no AI governance policy in place when it happened. Buyers have already priced this in. Madrona's 2026 Intelligent Applications 40 reports that data security and privacy is the number one criterion enterprises buy AI on, ranking top three for 78 percent of them, and that trust and governance has become a funded category of its own rather than a feature bolted onto something else. Security has stopped being the objection that delays the purchase. It is increasingly what is being purchased. Which tells you what a ban actually accomplishes. Usage stays right where it was and moves to a phone, under a personal login, where nothing is logged and nobody will mention it until something goes wrong. The tools your people use are simply ahead of the policies that cover them. The choice in front of you is between AI you can see and AI you cannot. ## I have been through the audits At Yembo I took an AI platform through ISO 27001, SOC 2 Type II, GDPR, and NIST 800-171 while continuing to ship. So I am not going to give your board a talk about how AI is risky, which they know, or a talk about how to slow down, which they will ignore. What I can give them is the distinction that matters: which controls genuinely reduce exposure, which ones are theater that makes an auditor comfortable and protects nothing, and what it costs to retrofit the real ones after a program has already shipped. Your security team and your AI ambitions are currently arguing, and in my experience both of them are right. For a working session rather than a talk, the 60-Minute Security Audit walks a leadership team through its own exposure and produces a prioritized remediation list. Teams in regulated industries usually pair it with the insurance governance material. ## What your audience walks out with - An honest estimate of how much unapproved AI use is already happening inside their organization. - The controls that genuinely reduce exposure, separated from the ones that only look like they do. - A sanctioned-tooling approach good enough that people stop reaching for personal accounts. - What ISO 27001, SOC 2 Type II, and GDPR actually ask of an AI system, from someone who answered it. ## Where I have taught and spoken ## Bring this to your board Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak about AI security and governance? Yes. Zach Rattner speaks to executive teams, security leaders, and boards on deploying AI without creating data exposure. He has taken an AI platform through ISO 27001, SOC 2 Type II, GDPR, and NIST 800-171, so the material comes from passing audits rather than from advising on them. What is shadow AI and why does it matter? Shadow AI is employee use of AI tools that an organization has not approved or cannot see. A Salesforce survey of more than 14,000 workers found that over half of those using generative AI at work use unapproved tools, and the IBM Cost of a Data Breach Report found shadow AI involved in one in five breaches, adding an average of $670,000 to the cost of each. Zach Rattner argues that a ban simply moves this usage onto personal accounts where nothing is logged. How should a company let employees use AI safely? Zach Rattner recommends providing a sanctioned tool good enough that people prefer it, combined with audit logging and clear data handling rules, rather than attempting prohibition. IBM found that 63 percent of organizations hit by shadow AI breaches had no AI governance policy, and 97 percent of those with AI-related breaches lacked proper AI access controls, which are exactly the controls a sanctioned tool provides and a ban does not. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run a security workshop? Yes. The 60-Minute Security Audit is a workshop for leadership teams of 8 to 25 people that walks an organization through its own AI data exposure, covering ISO 27001, SOC 2 Type II, and GDPR considerations, and produces a prioritized remediation list. --- ## Page: AI for fine art and museum logistics **URL:** /ai-for-fine-art-logistics **Description:** Keynotes on AI in fine art and museum logistics, from the CTO who built phone-based measurement-grade capture for galleries, vaults, and artwork. # AI for fine art logistics The job here is having the record you wish you had taken, before the object moved. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Museums & Shippers ## Fees from $10,000+ Book Now ## The claim is decided before the crate closes Ask your underwriter where your losses come from and the answer is the road. AXA XL's fine art underwriters say most of the claims they see are transit or water damage. Storage and display barely register. The damage happens once an object is in someone else's hands, in a truck, through customs, across a handoff between custodians your registrar has never met. The list of ways that goes wrong is longer than most people outside this work expect. Carriers writing inland marine cover for fine art underwrite against improper crating, shock and vibration, climate swings, water intrusion, and storage in uncontrolled conditions during a customs delay. Most of those leave no mark until someone opens the crate. And when a piece arrives damaged, the argument is never really about the damage. It is about what you can prove about the condition it left in. The claim is settled by the record built beforehand, which is careful, unhurried work, and it is usually the first thing squeezed when a show has to ship on a date. Everywhere else in logistics, the AI pitch is speed. In your world that pitch lands as slightly alarming. Nobody on your team wants a faster handler. Your failure mode is an incomplete record. Nobody has ever lost a claim over a slow one. ## What the technology already does today At Yembo I built ScanMyHome. Walk a gallery with a phone and what comes back is measurement-grade: the room's dimensions, the wall cutouts, the doorway you have to get a crate through, and the objects in the space identified individually. No tripod, no laser, no specialist. Do the same in your storage and your vault is mapped. Do it at a venue before a traveling show and you know what fits before the truck is loaded. A gallery is an interior space. So is a vault, a crate room, and the venue you are installing into next month. The geometry does not care what is hanging on the wall. None of that is a roadmap item or a demo. ScanMyHome's output exports into Xactimate, which means insurance carriers already settle property claims against those measurements, and that is a useful bar for whether a number is trustworthy enough to argue over. I am still not going to tell your registrars that software should write condition reports. It should not, and anyone selling you that has not handled anything fragile. A model that has seen ten million pieces of household furniture has never seen craquelure and has no business grading it. What it can do is guarantee that every object in a shipment was measured, photographed, catalogued, and timestamped, so your trained eyes spend their hours on the pieces that need judgment rather than the ones that only needed recording. The practical guidance on insuring art on the move comes back to the same thing: document before it ships, because nothing you do afterward substitutes. So what I bring your room is the view from inside a capability that already works on their spaces: what it does today, what it will be marketed as doing next year, and exactly where those two diverge. That is available about three years before the vendors selling into this market will admit any of it. ## The other side of the crate There is a second reason none of this is abstract to me. I make art. Lightwall is a kinetic light installation I built with the artist Rita Sus: glass, motors, radar, and a model that responds in real time to whoever walks up to it. It premiered at the California Center for the Arts Museum, and it is being prepared to exhibit at Dutch Design Week in Eindhoven in October 2026. Which puts me, this year, on the other end of your paperwork. Getting an installation of glass and embedded electronics from a museum in California to a venue in the Netherlands is a fine art shipping problem, and being the person whose work is inside the crate teaches you something that reading claims data does not. The record I would want if that piece arrives wrong is the record I am arguing for on the rest of this page. My work on AI in the moving industry covers the operational side of this in more depth, and my work on AI for insurance covers what happens on the carrier's side when a claim is contested. ## What your audience walks out with - A clear line between the documentation work a machine should carry and the judgment that stays with a trained registrar. - What a defensible visual record looks like when a claim is contested, and where most records fall short. - How to hold documentation standards across a partner network and custodians you do not employ. - What a documented inventory looks like when it is produced in minutes rather than budgeted as a week of registrar time. ## Where I have taught and spoken ## Bring this to your institutions Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak about AI in fine art and museum logistics? Yes. Zach Rattner speaks to fine art shippers, museum registrars, gallery operators, and the specialty moving companies that handle high-value collections. His talks cover AI-assisted condition reporting, inventory capture, and what a defensible record looks like when a claim is contested. How does AI help with fine art condition reporting? Roughly 60 percent of fine art insurance claims arise from transit, and those claims are settled on the documentation captured before an object moved. Zach Rattner argues that the useful role for AI in this work is producing a complete, timestamped visual record of every object before it moves. Claims are lost on incomplete records, and nobody in this industry is asking for faster handlers. What does Zach Rattner know about high-value logistics? Zach Rattner is the co-founder and Chief Technology Officer of Yembo, an AI platform used in more than 20 countries that turns a video walkthrough into a documented inventory with item-level detail. He holds 31 granted US patents, many covering computer vision applied to identifying and measuring objects in interior spaces. Can AI measure and inventory a fine art collection? Yes, and the technology is in production rather than experimental. ScanMyHome, which Zach Rattner built at Yembo, turns a phone walkthrough of an interior into a measurement-grade 3D reconstruction with floor plan, dimensions, wall cutouts, and individually identified objects. Insurance carriers settle property claims on that output, which is a useful benchmark for whether measurements are trustworthy enough to argue over. Can AI measure a gallery, a museum, or an individual artwork? Yes. ScanMyHome, built by Zach Rattner at Yembo, produces measurement-grade 3D reconstructions of interior spaces from a phone walkthrough, including floor plans, dimensions, wall cutouts, and identification of individual objects in the room. The underlying geometry problem is the same whether the space is a house, a gallery, a storage vault, or an installation venue, and the same capture measures the works in the room as well as the room itself. Has Zach Rattner exhibited work in a museum? Yes. Lightwall is a kinetic light and real-time AI installation Zach Rattner built with the artist Rita Sus. It premiered at the California Center for the Arts Museum and is being prepared to exhibit at Dutch Design Week in Eindhoven in October 2026. Zach Rattner is the technical director on the piece, which puts him on the institutional side of exhibition planning, installation, and international transit as well as the technology side. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run workshops for collections and logistics teams? Yes. Alongside keynotes, Zach Rattner runs working sessions for teams of 8 to 25 people, where the group maps its own condition reporting, inventory, and handoff workflows and leaves with a ranked list of where automation strengthens the record and where a trained human eye stays in the loop. --- ## Page: AI for home services: keynotes and workshops **URL:** /ai-for-home-services **Description:** Keynotes and workshops on AI for home services and the skilled trades, from the CTO who builds AI for field teams working in 20+ countries. # AI for home services You cannot hire your way out of a 110,000 technician shortage. You can stop spending the technicians you have on paperwork. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Operators & Franchises ## Fees from $10,000+ Book Now ## The labor math decides what AI is actually for The trades are short roughly 110,000 licensed technicians in 2026, and about half of employers report they cannot find skilled applicants. That gap does not close by recruiting harder, because a licensed tech takes years of apprenticeship to produce. The pipeline is the constraint. Which is part of why private equity is buying your competitors and, frankly, buying them for their people. When a licensed technician is the scarcest asset in the trade, acquiring a company becomes a hiring strategy. Construction is living the same arithmetic and needs 349,000 net new workers this year. So when a vendor pitches AI to your team as a way to run leaner, they have misread the market. Nobody in this trade is trying to need fewer technicians. You could not hire more if you wanted to. Every hour a licensed tech spends on paperwork, driving, or a callback is an hour of the scarcest resource in your industry spent on something unlicensed. ## What I bring to the room At Yembo I built AI that field teams in more than 20 countries use to assess a property from video, and produce the scope and inventory without sending someone to look first. That is the same problem your dispatch board has every morning, in a different uniform. I talk about what worked, what did not, and the part most speakers skip: how to tell a crew that software is going to touch their job without losing the best of them to the shop down the road. Technicians have heard the word automation before, and they did not hear it as good news. If your team wants to go deeper than a keynote, the agentic workflows workshop takes operations leaders through their own processes and ranks them by where automation pays back first. Contractors doing remodel and build work usually want to read about AI for construction as well. ## What your audience walks out with - A count of how many licensed hours a week are going to work that does not require a license. - Which parts of dispatch, estimating, and documentation are safe to automate now, and which need a person for good reasons. - Language for talking to crews about automation that does not read as a layoff announcement. - What a first deployment costs, measured against the fully loaded hourly cost of the licensed techs it gives back. ## Where I have taught and spoken ## Bring this to your operators Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak to home services companies? Yes. Zach Rattner speaks to home services operators, franchise networks, and trade associations across HVAC, plumbing, electrical, restoration, and field service. His keynotes cover where AI removes administrative load from licensed technicians and where it should stay out of the way. Will AI replace skilled trades technicians? Zach Rattner argues that it will not, because the industry cannot staff the roles it already has. Trade groups project a shortage of roughly 110,000 licensed technicians in 2026, and private equity buyers are acquiring contractors partly to obtain their licensed staff. In that market the useful role for AI is removing scheduling, documentation, and callback overhead so licensed hours go to licensed work. What experience does Zach Rattner have with field service work? Zach Rattner is the co-founder and Chief Technology Officer of Yembo, an AI platform used in more than 20 countries that surveys homes from video and produces inventories and estimates for field teams. He holds 31 granted US patents, many covering computer vision applied to physical spaces. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run workshops for home services leadership teams? Yes. Alongside keynotes, Zach Rattner runs working sessions for leadership teams of 8 to 25 people, where the group maps its own dispatch, estimating, and documentation workflows and leaves with a ranked list of what to automate first. --- ## Page: AI for insurance: keynotes and workshops **URL:** /ai-for-insurance **Description:** Keynotes and workshops on AI in insurance claims and underwriting, from a CTO who shipped computer vision into production and passed the audits. # AI for insurance Compliance is the only thing that will let you ship your AI program. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Carriers & Claims ## Fees from $10,000+ Book Now ## The regulator already arrived The NAIC Model Bulletin on the use of AI has now been adopted by roughly half the states, and the NAIC publishes a map of exactly which ones. If you write in more than a couple of jurisdictions, you are already inside it, whether or not your program has caught up. Transparency, explainability, and non-discrimination are no longer positions a carrier takes. They are things a carrier gets asked to demonstrate. Most claims organizations I talk to treat that as the brake. Legal is the reason the pilot has not shipped, the vendor's model is a black box, and the program is waiting on a governance review that keeps getting rescheduled. I think that reads your situation exactly backwards. Every carrier faces the same bulletin. The ones moving fastest right now are the ones who can answer the question when it arrives, without pausing the program to go find out. Explainability is the thing that lets you deploy at all. ## Why I can make that argument At Yembo I built computer vision that assesses property from video and put it into production claims and inspection workflows in more than 20 countries. I also took that platform through ISO 27001, SOC 2 Type II, GDPR, and NIST 800-171, so I have been on the receiving end of the questions your compliance team is about to ask, and I know which of them are real. That is the useful half of the talk. Not that AI can read a photo of a damaged roof, which your team already believes and your competitors already do, but which controls a claims model genuinely needs, which ones are theater, and what it costs to retrofit the difference after a program has already shipped. When your leadership wants to work rather than listen, the 60-Minute Security Audit takes them through the data-handling questions directly, my work on AI for customer service covers the claims contact center, and AI for fine art logistics covers the documentation side of a contested high-value claim. ## What your audience walks out with - A plain reading of what the NAIC bulletin actually requires of a claims model, separated from what vendors say it requires. - Where computer vision genuinely cuts claims cycle time, and where it quietly costs more than the manual step it replaced. - The controls worth building before a model goes live, from someone who has passed the audits. - A realistic 90-day view of a first claims deployment, including where the governance review actually stalls it. ## Where I have taught and spoken ## Bring this to your carriers Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak at insurance industry events? Yes. Zach Rattner speaks to carriers, claims organizations, and insurance associations on AI in claims, underwriting, and risk assessment. He has shipped computer vision into production claims workflows and speaks from deployment experience rather than from vendor research. How does AI regulation affect insurance carriers in 2026? The NAIC Model Bulletin on the use of artificial intelligence has been adopted by 23 states and the District of Columbia as of the first quarter of 2026, and a twelve-state pilot of the NAIC AI Systems Evaluation Tool began in early 2026. Zach Rattner argues that carriers who built explainability and model governance in from the start are now the ones able to deploy quickly, because they can answer a regulator without pausing the program. What does Zach Rattner know about insurance claims automation? Zach Rattner is the co-founder and Chief Technology Officer of Yembo, an AI platform used in more than 20 countries that assesses property from video for claims and inspection workflows. He holds 31 granted US patents, many covering computer vision applied to measuring and documenting physical spaces, and he has taken an AI platform through ISO 27001, SOC 2 Type II, GDPR, and NIST 800-171. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run workshops for insurance leadership teams? Yes. Alongside keynotes, Zach Rattner runs working sessions for leadership teams of 8 to 25 people. For insurance audiences these usually focus on the governance and data handling questions that decide whether a claims AI program can ship, and on ranking claims workflows by where automation pays back first. --- ## Page: AI for interior design: keynotes and workshops **URL:** /ai-for-interior-design **Description:** Keynotes and workshops on AI in interior design, covering room capture, visualization, and what automation does to a firm that bills by the hour. # AI for interior design Your taste is safe. The hours you bill around it are the problem, and nobody is talking about that honestly. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Studios & Trade Shows ## Fees from $10,000+ Book Now ## The threat is to the invoice, not the craft Showing your client three design directions used to mean briefing a studio and waiting somewhere between five days and three weeks, which one industry analysis puts at 15 to 25 billable hours. Firms using AI visualization now report reclaiming 14 or more hours a week. Every other industry hears that as pure upside. You should not, and I suspect you already know why. If you bill by the hour, those reclaimed hours are revenue that left the invoice. Your judgment is still in the work, the client still gets a better result, and the studio gets paid less for producing it. That is the actual mechanic behind the anxiety in this industry, and the tool vendors are never going to raise it. AI threatens a pricing model that charges for the parts a machine now does in thirty seconds. Your judgment stays in the work. ## Where I come into this At Yembo I built ScanMyHome: walk a room with a phone and get back a measurement-grade floor plan, dimensions, wall cutouts, and an itemized list of what is in it. Teams in more than 20 countries use that capture, and your first visit stops being a measuring exercise. That is the unglamorous half of a first visit, and it is the half genuinely about to change. I should be equally direct about what I am not. I am a technologist rather than a designer, I have never specified a room in my life, and nobody should book me to talk about taste. What I can tell you is which parts of your process a machine is genuinely good at now, which parts it is confidently bad at while sounding certain, and where that line sits this year rather than last. The other thing I bring is not a design subject at all. I have watched a professional services business reprice itself when software absorbed the hours it used to bill, which is the conversation your studio is going to have whether or not anyone schedules it. That transition is survivable, and it is far more survivable early. My work on AI for construction and remodeling covers the same capture problem on the build side of a project, which is usually the other half of this conversation. For studios weighing a working session instead of a talk, the agentic workflows workshop maps your own processes end to end. ## What your audience walks out with - An honest account of which parts of a design workflow AI does well today, and which it does badly while sounding confident. - The pricing conversation: how firms move from hours to value before the hours disappear. - What to tell clients who arrive with an AI-generated render and expect you to match it by Friday. - A realistic 90-day view of adopting these tools, including what it costs and where it usually stalls. ## Where I have taught and spoken ## Bring this to your studios Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs ## Questions program chairs ask Does Zach Rattner speak to interior design audiences? Yes. Zach Rattner speaks to design studios, industry associations, and the trade shows serving residential and commercial interior design. His talks cover how AI changes capture, visualization, and client iteration, and what that does to a firm that bills by the hour. Will AI replace interior designers? Zach Rattner argues that AI threatens the billing model built around drafting and visualization time, while design judgment stays in the work. Studios report reclaiming meaningful hours a week as concept visualization compresses from days to seconds, and AI-generated concepts still require professional validation for structural feasibility and code. The firms that do well are the ones that move toward value-based pricing before those hours disappear from the invoice. What does Zach Rattner know about capturing interior spaces? Zach Rattner is the co-founder and Chief Technology Officer of Yembo, an AI platform used in more than 20 countries that turns a video walkthrough of a room into dimensions and an itemized inventory. He holds 31 granted US patents, many covering computer vision applied to measuring and identifying objects in interior spaces. What does Zach Rattner bring to interior design audiences? Zach Rattner is a technologist rather than a designer, and Yembo is a capture and inventory platform rather than a design tool. What he brings to design audiences is the engineering behind room capture, and the experience of watching a professional services business reprice itself when software absorbs hours it used to bill. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Does Zach Rattner run workshops for design studios? Yes. Alongside keynotes, Zach Rattner runs working sessions for teams of 8 to 25 people, where the group maps its own intake, measurement, visualization, and client approval workflows and leaves with a ranked list of what to automate and what to keep as billable expertise. --- ## Page: AI for the moving industry: keynotes and workshops **URL:** /ai-for-moving-companies **Description:** Keynotes and workshops on AI for moving and relocation companies, from the CTO who built the video survey platform movers use in 20+ countries. # AI for the moving industry Your close rate is your growth problem, and your estimators are the ones who can fix it. I built the video survey AI that movers in 20+ countries run every day, and I speak about what it actually changed. ## Format Keynote or Workshop ## Delivery In Person & Virtual ## Audience Associations & Carriers ## Fees from $10,000+ Book Now ## The industry is buying AI to solve the wrong problem In 2024, 11.8% of Americans moved. That is down from 12.1% the year before and it is the lowest rate in Census records going back to 1948. A decade ago it was around 14 percent. In the 1960s it was 20 percent. Your addressable market has been shrinking for most of your career. Meanwhile fuel is up, insurance premiums are up, and agents are consolidating into holding alliances. Every AI vendor walking into your association meeting is selling the same answer to that squeeze: fewer estimator hours per job. I think that is the small prize, and it is the one that frightens your crews. When the number of moves is at a 77-year low, the scarce resource is the lead. The company that survives this consolidation is the one that quotes first and closes more of what it already had. Your close rate is your growth problem, and your estimators are the ones who can fix it. ## Why I can say that with a straight face At Yembo I built AI that surveys a home from a video walkthrough. It identifies the items, estimates the volume, and produces the inventory in minutes rather than a scheduled visit. Movers in more than 20 countries run it every day, which means I did not model this argument. I watched it happen, at scale, for a decade. I also watched the parts that went badly, and those are usually the more useful half of the session: the pilots that stalled, the crews who did not trust the first version, and the estimates the model got wrong before we fixed them. Your members have heard the sales pitch already. They have not heard an operator tell them which parts were oversold and what it cost to find out. ## The split nobody is planning for Consumer mobility is at a record low while corporate demand moves the other way. Atlas Van Lines surveyed 549 relocation decision-makers for its 2026 Corporate Relocation Survey and found more companies increasing employee moves and budgets, alongside more employees turning those moves down over housing and family. That is two different businesses pulling apart, and it changes where automation pays. Household goods, commercial relocation, and international moving each break in a different place: survey capacity in peak season, consistency across a partner network you do not control, square footage and elevator access instead of a living room. Fine art and high-value collections sit at the far end, where the inventory is also the insurance record and the tolerance for a missed item is zero. I shape the examples around whichever of those rooms I am standing in. If your members handle high-value collections, my work on AI for fine art logistics goes deeper on documentation, and AI for insurance covers what your carrier is doing with the same technology. ## What your audience walks out with - A clear read on which parts of a move AI can survey, quote, and dispatch today, and which still need a person in the room. - The questions to ask a vendor before signing, from someone who has spent years on the other side of that call. - What changes in an estimator's job, and how to tell your crews about it without losing your best people. - The one question to ask a survey vendor that separates a working product from a demo reel. ## Where I have taught and spoken "Zach's insights were practical, engaging, and very well received by the attendees." FIDI 39 Club Osaka, Japan ## Questions program chairs ask Does Zach Rattner speak at moving and relocation industry events? Yes. Zach Rattner has spoken to moving and relocation audiences including the Southwest Movers Association and the FIDI 39 Club in Osaka, Japan. His keynotes for this industry cover how AI changes surveying, estimating, and dispatch, and what it means for the estimators and crews doing the work today. What does Zach Rattner know about the moving industry? Zach Rattner is the co-founder and Chief Technology Officer of Yembo, an AI platform that moving companies in more than 20 countries use to survey homes from a video walkthrough and produce an inventory and an estimate without an in-home visit. He holds 31 granted US patents, many covering computer vision applied to measuring interior spaces. What are Zach Rattner's speaking fees? Zach Rattner's speaking fees start at $10,000 USD plus travel. The final number depends on the event type, the location, and whether the booking includes a workshop or a follow-up session. Is AI going to replace moving estimators? Zach Rattner argues that it will not, and that framing AI as an estimator replacement is the industry pursuing the smaller prize. With the US mover rate at its lowest level in Census records going back to 1948, the scarce resource is the lead rather than labor capacity. AI that returns a quote in minutes raises close rate, which is the growth lever in a shrinking market, and the estimator moves from driving between houses to closing business. Can Zach Rattner tailor a keynote to a moving industry audience? Yes. Zach Rattner builds the examples around the specific audience rather than delivering a stock talk. A room of van line executives, a room of independent agents, and a room of international relocation managers are three different conversations, and the keynote is shaped to whichever one is in front of him. Does Zach Rattner run workshops for moving companies? Yes. Alongside keynotes, Zach Rattner runs working sessions for leadership teams of 8 to 25 people, where the group maps its own surveying, estimating, and dispatch workflows and leaves with a ranked list of what to automate first and what to leave alone. ## Bring this to your members Tell me the audience and the date, and I will tell you whether I am the right fit. If your event is better served by someone else, I will say so. Most teams book the full day: the keynote plus the AI Opportunity Audit, a four-hour working session on your own workflows, at $25,000 flat. Check Availability See All Speaking Programs --- ## Page: Best local LLMs to run coding agents on Apple silicon **URL:** /projects/ai-mac-cluster/coding-models **Description:** Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for local agentic coding on Apple silicon, plus the memory math that decides. # Best local LLMs for agentic coding on Apple silicon You've decided to run your coding agent locally. Now comes the question that decides whether you'll actually like it: which model? Photo by Joey Banks on Unsplash. ## The hardware is a purchase. The model is a download. That asymmetry is the whole reason this page exists. Pick the wrong Mac and you live with it for years. Pick the wrong model and you type one Ollama command and try again. So people relax and grab whatever tops this month's leaderboard. Then the agent forgets its own plan mid-task, or generates at a pace that makes you nostalgic for dial-up, and they conclude local coding doesn't work. The model was wrong for the machine, not wrong in general. This guide is the missing selection step. The agentic coding setup guide covers getting the winner running. This page is about picking the winner. ## The four questions that pick the model A model card is a dating profile. Everyone lists their best benchmarks, nobody mentions what they're like to live with. These four questions are what actually predicts the experience, in the order they eliminate candidates. ### 1. Does it fit in memory the GPU can address? The weights and the context window have to sit in unified memory at the same time, and macOS only lets the GPU address roughly two-thirds to three-quarters of what the spec sheet says. A 48GB Mac presents about 36GiB to Ollama. If the model doesn't fit, nothing else about it matters. Remember that agentic coding is the hungriest context workload there is. The agent stuffs your codebase, its plan, and every tool result into the window, so budget tens of gigabytes beyond the weights, not hundreds of megabytes. ### 2. Dense or mixture-of-experts? Token generation speed is approximately memory bandwidth divided by the bytes read per token. A dense model reads every weight for every token: an 18GB model on a 273GB/s Mac mini tops out around 15 tokens per second, and on a 614GB/s M5 Max around 34. A mixture-of-experts model activates only a few billion parameters per token, so the same division yields several times the speed on the same machine. This is why the MoE question matters more on cheaper hardware. On a Mac mini, MoE is often the difference between usable and not. On an Ultra-class Studio with 1.2TB/s of bandwidth, dense models are already fast and their per-token quality tends to be steadier. ### 3. Is the context window agent-sized? Ollama's own documentation puts the floor for coding agents at 64,000 tokens. A model that tops out below that will lose the plot mid-task no matter how well it writes code, and a model trained for 256K gives the agent room to hold a real codebase in its head. Check the trained context length, not the default your runtime happens to configure, which is a separate trap the setup guide covers in detail. ### 4. Coding-tuned or general? At equal size, a coding-tuned model usually writes better code. But agentic coding rewards more than code: the loop lives or dies on instruction following, tool calling, and staying on task across a long session, and strong general models are often better at exactly that. This is the question benchmarks answer worst, and the reason the final call belongs to a two-week trial on your own backlog rather than to a leaderboard. ## The contenders Five open-weight models are worth your download bandwidth right now. Sizes below are the quantized builds you would actually run, 4-bit unless noted. Model Architecture 4-Bit Size Trained Context Role Qwen3.8-27B 27B dense 18GB 256K The all-rounder this series runs Qwen3-Coder-30B-A3B 30B MoE, about 3B active 19GB 256K The speed pick for modest bandwidth Gemma 4 26B 26B dense, multi-token prediction 16GB 128K The efficiency pick DeepSeek V4 Flash 284B MoE, about 13B active 155GB, or 90GB at 2-bit 1M The flagship a Mac Studio can actually house GLM-5.3 753B MoE, about 40B active 425GB, heavy quants smaller 1M The open-weight ceiling, for 512GB Ultras ### Qwen3.8-27B: the default This is the model the rest of this series runs, and the reasons are the four questions above rather than sentiment. It fits comfortably at 64GB, its 256K trained context is genuinely agent-sized, and its instruction following is what makes agent mode viable at all. The qwen3.8 tag on Ollama is one command away. If you have the bandwidth to run a dense 27B at usable speed, start here. ### Qwen3-Coder-30B-A3B: the speed pick The name decodes as 30 billion parameters with about 3 billion active per token, and that second number is the one your memory bus feels. Reading a tenth of the weights per token moves the generation ceiling from teens to triple digits on the same Mac mini. It is also coding-tuned, so on constrained hardware you're not trading quality for speed so much as trading a little generality for a lot of throughput. Grab it from the qwen3-coder tag on Ollama. If your cluster is Mac minis rather than Studios, this is probably your model. ### Gemma 4 26B: the efficiency pick Google's entry ships with multi-token prediction, which lets speculative decoding draft several tokens per read and verify them in one pass. In practice that recovers a chunk of the MoE speed advantage while keeping a dense model's consistency. Its 128K context is half of what the Qwen pair trains for, which for agentic work on a large codebase is a real limit rather than a spec-sheet footnote. Pull the 26B from the gemma4 tag and keep it resident as the second model for review passes and quick edits. ### DeepSeek V4 Flash: the flagship a Studio can house This is the model that moved the local ceiling this year. 284 billion parameters with about 13 billion active per token means it reasons like a frontier model and generates like a mid-size one, and its 1M-token context window is trained in rather than bolted on. Community testing puts the 4-bit build around 35 tokens per second on a 512GB M3 Ultra, which is a frontier-class agent running at usable speed with the meter off. The catch is residency. The 4-bit conversions land around 155GB, which wants a 256GB Studio, and the aggressive 2-bit builds near 90GB squeeze onto a 128GB machine with a real quality tax. If you specified an M5 Ultra from the hardware comparison, this is what all that unified memory is for. The deepseek-v4-flash tag carries it. ### GLM-5.3: the open-weight ceiling Z.ai's flagship is the most capable open-weight coding model you can download right now, built specifically for coding agents and scoring 88.2 on Terminal-Bench 2.1. At 753 billion parameters with about 40 billion active, even the 4-bit build is around 425GB, so a genuinely local run means a 512GB Ultra-class Studio and an aggressive quant, and most teams will run it hosted instead. Its smaller multimodal sibling on the glm-5.3 family tags, GLM-5.3-Flash with 18 billion active parameters, is the one to watch for the local tier. ### The ones you license instead of house Open weights and locally runnable are different claims, and the gap between them is now hundreds of gigabytes. Kimi K3 is a 2.8-trillion parameter monster whose 1-bit quant still wants a 650GB floor, per Unsloth's deployment docs. DeepSeek V4 Pro at 1.6 trillion and the 2.4-trillion Qwen3.8 flagship are the same story. MiniMax-M3 at 428 billion is the borderline case, quantizable onto a 512GB Ultra if you want its multimodal input. For everything in this paragraph, you are choosing a hosting provider, not a Mac, and that decision has its own compliance math covered in the security and compliance article. ## Match the model to the Mac Cross the four questions with the hardware you own and the field usually collapses to one obvious answer per machine. Your Mac Run This Why 32GB Mac mini Qwen3-Coder-30B-A3B MoE speed makes modest bandwidth usable, and the context window has to stay capped anyway at this memory size. 64GB Mac mini or MacBook Pro Qwen3.8-27B The series pick. Full 256K context fits, and the dense model's steadiness pays for itself across long agent sessions. 128GB M5 Max Mac Studio DeepSeek V4 Flash at 2-bit The 90GB builds put a 284B-class agent on a desktop. Tight on context, and the quality tax is real, so weigh it against Qwen3.8-27B with room to breathe. 256GB Ultra-class Mac Studio DeepSeek V4 Flash at 4-bit The 155GB conversion fits with agent-sized context to spare, at roughly 35 tokens per second on Ultra-class bandwidth. 512GB M5 Ultra Mac Studio GLM-5.3, heavily quantized The only Mac where the open-weight ceiling is reachable at all, with DeepSeek V4 Flash at full quality as the safer resident. To check any pairing not listed here, the local LLM memory and speed calculator estimates required unified memory and generation speed for every chip from M1 to M6. And if the answer to the hardware question is still open, the M5 Ultra vs. M5 Pro vs. M6 comparison is the buying decision this table assumes you've already made. Whichever row is yours, the final step is the same: pull the model, set the context window correctly, and wire it into your editor. That is exactly what the agentic coding setup guide walks through, including the context window trap that makes a right model feel like a wrong one. ## Frequently asked questions ### What is the best local LLM for agentic coding on a Mac? For most Macs with 64GB of unified memory or more, Qwen3.8-27B is the strongest all-round choice: a dense 27-billion parameter model with a 256K context window that follows agent instructions reliably. On machines with less memory bandwidth, such as a Mac mini, Qwen3-Coder-30B-A3B is the better pick, because its mixture-of-experts design reads only about 3 billion parameters per token and generates several times faster on the same hardware. ### Is a mixture-of-experts or dense model better for local coding? It depends on memory bandwidth. Token generation speed is roughly memory bandwidth divided by the bytes read per token. A dense model reads all of its weights for every token, so an 18GB model on a 273GB/s Mac mini tops out around 15 tokens per second. A mixture-of-experts model activates only a fraction of its weights per token, so it can generate several times faster on the same machine. On high-bandwidth hardware like an Ultra-class Mac Studio, dense models close the gap and their per-token quality tends to be more consistent. ### Can you run DeepSeek V4 locally on a Mac? DeepSeek V4 Flash runs locally on a Mac Studio. It is a 284-billion parameter mixture-of-experts model with about 13 billion parameters active per token and a 1M-token context window. The 4-bit conversions land around 155GB, which wants a 256GB or 512GB Studio, and the aggressive 2-bit builds near 90GB fit a 128GB machine at reduced quality. Community testing reports roughly 35 tokens per second on a 512GB M3 Ultra. The larger DeepSeek V4 Pro is API-tier hardware and does not fit any single Mac. ### What is the most capable open-weight model for coding? GLM-5.3 from Z.ai is the most capable open-weight coding model as of late 2026, scoring 88.2 on Terminal-Bench 2.1 and built specifically for coding agents. It is a 753-billion parameter mixture-of-experts model with about 40 billion active per token and a 1M context window, so running it locally means aggressive quantization on a 512GB Ultra-class Mac Studio. Most teams run GLM-5.3 hosted and keep a smaller model like Qwen3.8-27B or DeepSeek V4 Flash local. ### How much unified memory do you need for a local coding model? 32GB is the practical floor and runs a 27B-class model with a constrained context window. 64GB is where local agentic coding stops feeling like a compromise, because macOS only lets the GPU address roughly two-thirds to three-quarters of unified memory, and an agent needs tens of gigabytes for context on top of the weights. 128GB opens up the 2-bit builds of DeepSeek V4 Flash, 256GB runs its 4-bit conversions, and 512GB is where the GLM-5.3 class becomes possible at all. ### Do coding-tuned models beat general models for agentic work? At the same size and speed, a coding-tuned model like Qwen3-Coder usually wins on code generation, but agentic coding rewards more than code quality. The agent loop depends on instruction following, tool calling, and staying on task across a long session, and strong general models like Qwen3.8-27B are often better at exactly that. The honest answer is to run the two-week test: point each candidate at the same real backlog tasks and keep the one that finishes more of them. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 Currently Reading ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Current Page Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Picking models is the easy half Getting an engineering org to actually adopt local AI is the hard half, and it is what I speak about. My hands-on workshop, Modernize Your Engineering Org for the AI Era, takes teams from their first local model to a production agentic coding stack. See AI Speaking Programs Book the Agentic Coding Workshop --- ## Page: M5 MacBook Air LLM benchmark: sustained AI performance **URL:** /projects/ai-mac-cluster/fanless-benchmarking **Description:** Six ten-minute runs on a fanless M5 MacBook Air. Sustained throughput holds 69 to 77 percent of cold, and plugging in made it slower. # Can the fanless M5 MacBook Air handle heavy AI workloads? The MacBook Air is popular because of its long battery life, silent operation, and sleek design. But how much does that packaging cost you in terms of running local LLMs? Photo by Yudhajit Ghosh on Unsplash. Here is a thing that happens to everyone who runs a model locally for the first time. You read a benchmark, you buy the machine, you load the weights, and the first few requests are gorgeous. Then you point an agent at a real repository, walk away to get coffee, and come back to something noticeably worse. You didn't do anything wrong. Neither did the benchmark. It just stopped watching too early. Every benchmark I've published in this series, including the fifty-four runs comparing MLX against Ollama, reports a median of three repetitions. That's the right summary for a quantity that holds still. On a laptop with no fan it doesn't hold still — and a median is the statistic that hides it. Three repetitions finish inside the first ninety seconds. On this machine that whole window sits above the rate it can hold. So I rebuilt the harness to report throughput over time instead of a single number, and ran it for ten minutes a cell. ## M5 MacBook Air AI performance: plan on 70 to 75 percent of the benchmark That's the boring conclusion. The more interesting one is what sits underneath the cliff. There's a class of work that finishes before the machine ever settles, and another that idles long enough between requests to shed the heat it made, and neither of those is redlining anything. I expect the good fanless use cases to come from that shape rather than from anyone squeezing a batch job onto a laptop. If you're going to run it anyway, and I do, here is how to get the most out of the hardware. - Budget on 70 to 75 percent of any published cold number. Two models of different sizes and architectures, both phases of inference, and both power sources all landed between 69 and 77 percent. - You are off the cold rate before a three-run benchmark finishes. Every cell dropped below 95 percent of its opening bucket within 30 to 60 seconds. - It settles, it does not spiral. The decline finishes at roughly 150 seconds and then holds flat for the rest of the run, so there is a steady state you can plan against. - Waiting for the first token gets worse faster than tokens per second does. A 4,096 token prompt went from 3.74 to 5.10 seconds on the 30B, and from 11.1 to 14.4 seconds on the dense 14B. - Unplugging made it faster. Mains was 7 to 12 percent slower than battery on every sustained figure, and the two are indistinguishable while the machine is still cold. - Continuous inference costs about a quarter of the battery per hour, so a full charge is worth roughly three and a half hours of it. ## Qwen3 tokens per second on the M5 MacBook Air, cold against sustained Six cells, 487 individual measurements, zero failures. Retention is the last thirty-second bucket divided by the first, so a machine that holds its rate reports 100 percent. ### Sustained throughput on a fanless M5 MacBook Air #### Prompt processing tokens per second, 4,096 token prompt, higher is better #### Token generation tokens per second, 256 generated tokens, higher is better Cold, first 30 seconds Sustained, after 10 minutes Higher is better zachrattner.com Cold is the first thirty-second bucket and sustained is the last, from ten-minute runs. The two panels carry separate scales because the two phases of inference differ by more than an order of magnitude. Prompt processing, tokens per second, 4,096 token prompt. Higher is better. Cell Cold Sustained Retention Qwen3-14B 4-bit, battery 355 274 77.2% Qwen3-14B 4-bit, mains 359 254 70.7% Qwen3-30B-A3B 4-bit, battery 1063 788 74.1% Qwen3-30B-A3B 4-bit, mains 1040 725 69.7% Token generation, tokens per second, 256 generated tokens. Higher is better. Cell Cold Sustained Retention Qwen3-30B-A3B 4-bit, battery 57 44 76.9% Qwen3-30B-A3B 4-bit, mains 57 39 68.6% All figures are tokens per second. Prefill is prompt processing, measured at a 4,096 token prompt. Decode is token generation, measured at a 512 token prompt and 256 generated tokens. The interesting part isn't the size of the drop. It's the consistency. Two models of different sizes and architectures, both phases of inference, both power sources, all land between 69 and 77 percent. That's a property of the chassis rather than of any model — so you can apply it to a number somebody else published. ## Time to first token on the M5 MacBook Air rises 30 to 41 percent Nobody experiences tokens per second. What you feel is the gap between hitting enter and seeing the first word, so the harness records that per bucket too. ### Time to first token on a fanless M5 MacBook Air #### 4,096 token prompt seconds to first token, lower is better #### 512 token prompt seconds to first token, lower is better Cold, first 30 seconds Sustained, after 10 minutes Lower is better zachrattner.com Seconds from sending the prompt to receiving the first token, cold against sustained. Prompt length is the dominant term in this measure, so the two prompt sizes are plotted on separate scales rather than crushed onto one. 4,096 token prompt, seconds to first token. Lower is better. Cell Cold Sustained Longer by Qwen3-30B-A3B, battery 3.74s 5.10s 36% Qwen3-30B-A3B, mains 3.87s 5.47s 41% Qwen3-14B, battery 11.10s 14.40s 30% Qwen3-14B, mains 11.00s 15.50s 41% 512 token prompt, seconds to first token. Lower is better. Cell Cold Sustained Longer by Qwen3-30B-A3B, battery 0.68s 0.93s 38% Two things fall out of that. The percentages are bigger than the throughput percentages, because latency is the reciprocal of a rate and a 25 percent drop in rate is a 33 percent rise in wait. And the dense 14B is three times slower in absolute terms despite being less than half the size. That's the mixture-of-experts architecture doing its job. The 30B activates about 3B parameters per token. The 14B activates all fourteen, so the smaller model does more arithmetic per token than the bigger one. The same decline shows up in the plainest unit there is — how much work finishes. In the first minute of the 30B run on battery the machine completed 15 requests against 4,096 token prompts. In the last minute it completed 11. ## M5 MacBook Air thermal throttling settles after about 150 seconds This matters more than the headline number, because it changes what you do about it. The machine isn't degrading without limit. It finds a rate it can hold, and then it holds it. A runner who opens a marathon at sprint pace isn't cheating and isn't broken. They're going to settle into the pace the body can hold. The only mistake available is timing them over the first two hundred meters and calling that the marathon time. ### Throughput across a ten-minute run, battery against mains #### Prompt processing tokens per second, higher is better #### Token generation tokens per second, higher is better On battery On mains Higher is better zachrattner.com Every thirty-second bucket of the two 30B cells. The fall happens in two stages and is over by about 150 seconds, after which both curves run flat. Battery and mains open together and separate only once the chassis has warmed up. Prompt processing, tokens per second, per thirty-second bucket. Higher is better. Elapsed seconds On battery On mains 0 1063 1040 30 996 990 60 983 945 90 928 886 120 872 757 150 790 721 180 769 716 210 787 708 240 774 705 270 791 705 300 807 729 330 782 707 360 808 733 390 805 733 420 798 723 450 806 728 480 794 739 510 779 739 540 802 722 570 788 725 Token generation, tokens per second, per thirty-second bucket. Higher is better. Elapsed seconds On battery On mains 0 57 57 30 54 54 60 52 49 90 45 41 120 43 36 150 40 35 180 41 36 210 42 37 240 42 38 270 42 38 300 43 39 330 44 38 360 44 38 390 44 38 420 43 39 450 44 39 480 43 39 510 43 39 540 44 39 570 44 39 600 44 not recorded Both curves hit their floor at about 150 seconds and stay there for the remaining seven and a half minutes. Decode even recovers a little, from a trough of 40.0 to 44.1 in the final bucket. The control model throttles in two distinct stages: a settle in the first thirty seconds, a plateau near 330 tokens per second for two and a half minutes, a second step down at around 180 seconds, and then flat at 274 for the rest of the run. A benchmark that stops at ninety seconds sees only the first stage, and publishes the plateau as though it were the answer. ## Thermal throttling, not memory pressure, on a 32GB M5 MacBook Air This is the first objection any competent engineer raises, and it's the right one. A 30B model at 4-bit is roughly 16GB on a 32GB machine. Any slowdown might be the system paging rather than the chip getting hot — and if it were, everything above would be a story about memory wearing a thermal costume. That objection is why the 14B model is measured at all. Qwen3-14B at 4-bit is 8.3GB. It's dense rather than sparse, so it does more arithmetic per token than the mixture-of-experts model and generates more heat, not less. Across all twenty of its buckets, swap didn't move by a single megabyte and free memory never fell below 54 percent. It lost 23 percent anyway. The decline is thermal. Memory pressure is a real effect and it's a different one. It changes the level rather than the slope. An earlier run with 9GB of swap in use opened at 1,035 tokens per second against 1,131 on the clean machine, about 9 percent lower, and then declined at the same rate. Both are true, and they're worth keeping apart when you're working out why your own machine is slow. ## Heat is the enemy: the M5 MacBook Air ran slower plugged in than on battery Mains power was slower than battery in every single cell, by 7 to 12 percent on the sustained figure. Cold, the two are indistinguishable. The 14B opens at 355.4 on battery and 359.3 on mains, and the curves track each other within a couple of percent for the first three and a half minutes. Then the battery run flattens at 275 and the mains run keeps sliding to 254. They start together and diverge only as heat accumulates. That points at the charging circuitry putting heat into a chassis with no fan to get it out. It's the mechanism I'd guess at, and I want to be clear that it's a guess. What's measured is the divergence. The cause is not. The conditions are worth being precise about, because an earlier attempt at this comparison was ruined by them. The mains run started with the battery at 99 percent, reporting finishing charge with zero minutes remaining, after the machine had sat idle for about thirty-four minutes. So this isn't a battery being charged from empty. It's a full battery on a cold machine. ## Local LLM response times on the M5 MacBook Air run about a third longer Total time for a request is the prompt divided by the prefill rate plus the output divided by the decode rate. So the honest way to express the gap is to price a few real tasks at both rates. These use the 30B on battery, the best case measured. ### What one request costs at cold and sustained rates #### One request, end to end seconds, lower is better Cold, first 30 seconds Sustained, after 10 minutes Lower is better zachrattner.com Seconds to finish one request, priced from the 30B on battery at both rates. Time equals the prompt divided by the prefill rate plus the output divided by the decode rate. One request, end to end, seconds. Lower is better. Cell Cold Sustained Longer by Agentic coding edit 12.8s 16.9s 33% Summarize a long document 23.8s 31.6s 33% Chat turn 14.4s 18.8s 30% Short question 2.9s 3.8s 30% The penalty is close to uniform at about a third, because both phases throttle by similar proportions. That's a convenient result. Take a published cold benchmark for this machine, add a third, and you'll be close. ## MacBook Air battery life running a local LLM: about three and a half hours The harness samples charge level per bucket, so the battery cells answer this for free. The ten-minute prompt processing run took the machine from 95 to 91 percent. The token generation run that followed took it from 91 to 86. That's 24 to 30 percentage points an hour, so a full charge buys somewhere around three and a half hours of back-to-back inference. Treat it as an extrapolation from ten-minute cells rather than a measured discharge — the last twenty percent of a battery doesn't behave like the first. It's still the right order of magnitude for deciding whether you can do this on a plane. ## How to work around thermal throttling on a fanless Mac The honest headline is the one you already suspected. A fanless laptop is the wrong machine for serious sustained AI work, and nothing measured here rescues it. What the numbers add is the where and the why: the cliff arrives inside the first minute, it bottoms out at about 150 seconds, and it costs you a third of your throughput and a quarter of your battery an hour. - Size batch jobs on the sustained rate, not the cold one. Anything running longer than about three minutes spends nearly all of its life in the flat part of the curve, so the sustained figure is the one to divide by. Estimate a two-hour job from a three-repetition benchmark and you'll be about a third short. - Interactive use never gets there. A chat turn is a few seconds of load followed by however long you spend reading, and the machine sheds heat in the gap. Duty cycle is why the same model feels fine in a chat window and disappointing in an agent loop — real enough that this harness had to move prompt construction outside the timed loop to stop the cooling gaps inflating its own results. - Unplug for long runs. Counter-intuitive, free, and worth 7 to 12 percent on this chassis. It costs about a quarter of the battery an hour to take. - Do not use the thermal pressure counter as your trigger. On the 14B cells it still read nominal 90 seconds in, by which point throughput was already down 7 percent. It tells you the system has decided to shed performance, not how much and not when it started. Throughput is the measurement. The counter is corroboration. - If sustained throughput is the job, buy a fan. Everything on this page is a property of a chassis with no way to move heat out of it. The hardware comparison in Part 2 is where that decision gets made, and a laptop is the wrong end of it for these kinds of workloads. ## Run this MLX benchmark on your own Mac Your chassis isn't this chassis, and the whole point of publishing the harness is that you don't have to take my retention figure for yours. It needs mlx-lm and nothing else. $ uv venv --python 3.14 .venv $ uv pip install --python .venv/bin/python -r requirements.txt $ .venv/bin/python bench-thermal.py --verify-cache $ .venv/bin/python bench-thermal.py --smoke $ .venv/bin/python bench-thermal.py --duration 600 Run the two short commands before the long one. The first sends the same prompt twice and then a fresh one, then prints whether the engine cached the prefix and whether the nonce defeated it. On this machine it reports no cache to defeat. That's the finding that lets everything else stand. The second pulls a 0.6B model and runs for sixty seconds, proving the plumbing without spending ten minutes to discover a model key was wrong. The harness refuses to start in Low Power Mode, because that caps performance and would be recorded as thermal throttling. It cools to nominal, runs a discarded warmup so kernel compilation is paid before the clock starts, cools again, and only then measures. Every raw result behind this page is in the gist. ## How to benchmark thermal throttling on Apple silicon MacBook Air with an M5 and 32GB, macOS 26.6.2, Python 3.14.7, mlx 0.32.2 and mlx-lm 0.31.3. Ten minutes per cell, thirty second buckets, median within each bucket. Four things had to be controlled, and every one of them would have quietly invalidated the run. Prompts are built before the clock starts. Constructing a 4,096 token prompt means tokenizing repeatedly to hit the target, which takes long enough for a fanless machine to shed heat between iterations. Doing that inside the loop would measure a duty cycle rather than sustained load, so a pool of sixteen prompts is built up front and the loop only generates. Kernel compilation is paid before the clock starts. The first call into MLX compiles, and that looks like a cold machine running fast and then settling. Measured on this machine it is worth 1.5 times, so the sequence is cool, one discarded warmup, cool again, then measure. Every prompt carries a nonce. Prompt caching is what produced the 474 times fiction documented in the previous article, and a cached prefill would fabricate these numbers outright. MLX turns out to have no cache to defeat, which was verified three times on this machine before anything else was trusted. Swap and free memory are recorded per bucket. Without that the memory question above could not be answered, and a decline caused by paging would have been published as a decline caused by heat. One defect worth disclosing, since the raw files are published. The charging field in the two battery result files reads true and is wrong, because the detector tested for a substring that also appears inside the word discharging. It's metadata, no measurement depends on it, and it's fixed in the harness. The data files are left as they were produced rather than edited after the fact. ## What this M5 benchmark does not measure: Mac Studio, M4, and active cooling Apple published its own MLX figures for the M5, reporting prompt processing 3.33 to 4.06 times faster than the M4 and generation 1.19 to 1.27 times faster, measured on a MacBook Pro at a 4,096 token prompt. Nothing on this page tests that. Apple compared two chip generations. This compares one chip against itself over ten minutes. The two are complementary, and it would be wrong to read either as contradicting the other. Nor does this page say anything about an actively cooled machine. A Mac Studio has a fan and I'd expect it to hold close to 100 percent. But expecting isn't measuring, and if it turned out to sag under sustained load that would be a larger finding than anything here. For the same reason I've deliberately not compared these numbers against the M1 Ultra figures in the MLX and Ollama article. Those are medians of three runs, and this page has just shown that overstates sustained throughput by about 20 percent on a machine like this one. Putting a sustained figure next to a burst figure and calling the difference a generational comparison would be the error this whole series exists to document, pointed at my own data. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. ## Frequently asked questions about local AI on the M5 MacBook Air ### Does the M5 MacBook Air throttle when running local LLMs? Yes, and it settles rather than collapsing. On an M5 MacBook Air with 32GB, sustained prompt processing fell to between 69 and 77 percent of its cold rate across six ten-minute runs. Every run was already below 95 percent of its cold rate within the first 30 to 60 seconds, the decline finishes at about 150 seconds, and the rate then holds flat for the remaining eight minutes rather than degrading without limit. ### How much slower is a MacBook Air after ten minutes of AI load? On a fanless M5 MacBook Air, about 23 to 31 percent slower than the first thirty seconds. A dense Qwen3-14B model at 4-bit went from 355 to 274 tokens per second of prompt processing on battery. A Qwen3-30B-A3B mixture-of-experts model went from 1,063 to 788. Token generation fell from 57.4 to 44.1 tokens per second over the same period. ### How much does thermal throttling add to time to first token on a Mac? Between 30 and 41 percent, measured on an M5 MacBook Air across six ten-minute runs. A 4,096 token prompt sent to a Qwen3-30B-A3B mixture-of-experts model returned its first token in 3.74 seconds cold and 5.10 seconds after ten minutes of continuous load. The same prompt sent to a dense Qwen3-14B model went from 11.1 seconds to 14.4 seconds. Time to first token degrades by a larger percentage than tokens per second does, because latency is the reciprocal of a rate. ### How long does an M5 MacBook Air battery last running a local LLM? Roughly three to four hours of continuous inference. On an M5 MacBook Air with 32GB, a ten-minute sustained prompt processing run took the battery from 95 to 91 percent, and the ten-minute token generation run that followed took it from 91 to 86 percent. That is 24 to 30 percentage points per hour, extrapolated from ten-minute cells rather than measured as a full discharge. ### Does running on battery slow down local AI inference on a Mac? No. On an M5 MacBook Air the opposite happened: mains power was 7 to 12 percent slower than battery under sustained load. Cold, the two are within 1 percent of each other. They diverge only as heat accumulates, which points at the charging circuitry adding heat to a chassis that has no fan to remove it. This was measured with the battery at 99 to 100 percent and not actively charging. ### Is 32GB enough to run a 30B model on an M5 MacBook Air? Yes. A Qwen3-30B-A3B mixture-of-experts model at 4-bit occupies roughly 16GB and ran on a 32GB M5 MacBook Air with free memory holding at 29 to 31 percent and no swapping during the measured runs. Memory pressure is a separate problem from thermal throttling: a run with 9GB of swap in use started about 9 percent slower but declined over time at the same rate. ### Why do local LLM benchmark numbers not match what I actually get? Because most benchmarks report a median of three runs, and three runs finish inside the first ninety seconds. On a fanless machine that window sits above the sustained rate. A three-repetition benchmark of a 14B model on an M5 MacBook Air would report roughly 330 tokens per second where the machine sustains 274, overstating it by 20 percent. ### How do you measure thermal throttling on a Mac? Run one workload continuously for a fixed duration and report the median inside each time bucket rather than a single median across the whole run. Record thermal pressure from com.apple.system.thermalpressurelevel alongside it, but treat throughput as the measurement and the pressure level as corroboration, because the operating system signal lags: on an M5 MacBook Air it still read nominal 90 seconds into a run where throughput had already dropped 7 percent. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Read Article Part 9 Currently Reading ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Current Page ### Measuring the right thing is the hard part A number that's wrong in a way nobody notices is worse than no number at all. That's true of a laptop benchmark, and it's true of every AI pilot that reported a win nobody could reproduce. If you're the one who gets handed those numbers, the workshop spends a day on how to check them. See the workshop See AI speaking programs --- ## Page: MLX vs Ollama on Apple silicon: measured benchmarks **URL:** /projects/ai-mac-cluster/mlx-vs-ollama **Description:** I measured MLX against Ollama on the same weights and matched quantization. MLX wins dense 4-bit by 41 percent, ties at 8-bit, and loses on MoE. # Is MLX faster than Ollama? It depends. It wins by 41 percent on one model. It loses by 22 percent on another. Same machine, same afternoon. Fifty-four runs on an M1 Ultra. Here is what actually decides it. I run a rack of Mac minis in production at Yembo. It replaced about $40,000 a year of cloud spend, though that particular workload is speech transcription running on Whisper, not a language model. The language models are a separate job on the same hardware, and they are the ones I use every day for agentic coding. So when people started saying Apple's MLX was much faster than Ollama, the question was narrow and practical: am I leaving speed on the table every time I open my editor? To get to the bottom of it, I decided to run a good old-fashioned bake off. ## The bake off Three matchups. In each one, both engines get the same model, converted from the same upstream weights, so the only thing changing is who is doing the math. - Qwen3.8-27B at 4-bit. The everyday case, and the model I actually code against. Ollama's build against the MLX one. - The same model at 8-bit. This is the control, and it is the row that ends up explaining everything else. MLX 8-bit against Ollama's Q8_0. - Qwen3-30B-A3B at 4-bit. A mixture-of-experts model, to see whether the answer holds when the architecture changes. The MLX build against the matching Ollama tag. Each one gets three prompt lengths, roughly 500, 4,000 and 16,000 tokens, because a chat message and a codebase are different problems. Each generates 256 tokens. Three runs per combination, engines alternating so neither one eats the other's heat. And two numbers come out separately rather than blended: how fast it reads your prompt, and how fast it writes the reply. Those are different problems with different limits, which is worth two minutes before the tables. Fifty-four runs in total. The first answer I got was wrong, and wrong in a way that would have been very easy to publish. I will come back to that. ## The short version On a dense 27B model at 4-bit, MLX generates 41 percent more tokens per second. Switch that same model to 8-bit and the lead vanishes. Switch to a mixture-of-experts model and Ollama wins by 22 percent. Same hardware, same weights, three different answers. Which means the interesting question is not which engine wins. It is why the answer keeps moving. ## 9/2/2026 update Reid Peryam read this and asked a good question in the comments: was I using Rapid-MLX, rather than vanilla MLX? I was not. So I ran it. A well-built wrapper buys a few percent over the library it wraps, and it changes nothing about which engine to pick. Rapid-MLX is a server built on the same MLX that this page benchmarks as a library. On the dense 27B at 4-bit it generates 1.9 to 4.4 percent more tokens per second than vanilla MLX. On the mixture-of-experts model, Ollama still wins, and by more than it wins against MLX alone: 20 to 22 percent ahead of Rapid-MLX at every prompt length. Its repository headline claims 4.2x faster than Ollama. I measured 1.40x to 1.53x on the dense model, and 0.79x on the mixture-of-experts one, which is to say slower. Rapid-MLX's own published benchmark measures 1.46x and 1.50x, so my numbers reproduce theirs almost exactly. The headline is the part nobody can support. And the prompt cache trap that nearly wrecked the original run happens here too, at 74x. ### What I tested Three engines instead of two, on the same afternoon, the same Mac Studio, and the same two checkpoints from the tables below. Fifty-four more runs. Rapid-MLX was served the identical MLX repository that the vanilla MLX column loads, so this is a wrapper measured against the library inside it rather than against a different model. One thing had to change, and it is the whole reason this took a day rather than an hour. Rapid-MLX exposes no timing counters of its own: token counts, and nothing else. Ollama and MLX both report their own nanosecond durations, which is what every number above is built from. Timing Rapid-MLX from the outside while reading the other two from the inside would have charged it for HTTP that vanilla MLX never pays, and that is the same shape of mistake as the cache artifact, a number that looks like a fair fight and is not one. So all three engines are timed from the client here, identically, and their own counters are recorded alongside for checking. The two agreed within about 3 percent throughout. Prompt token counts were matched across all three engines on every repetition, and the run fails loudly if they ever diverge. Decode below is the median of three runs. Model Prompt Ollama MLX Rapid-MLX 27B dense 512 20.86 30.72 31.86 27B dense 4,096 22.07 29.61 30.92 27B dense 16,384 19.00 27.41 27.95 30B-A3B MoE 512 91.23 68.06 72.50 30B-A3B MoE 4,096 80.50 60.01 64.00 30B-A3B MoE 16,384 57.76 45.62 45.01 The verdict further down this page said to stay on Ollama for mixture-of-experts models. That survives, and it comes out stronger: Ollama beats both MLX-based engines at every prompt length. It also survives the one discrepancy worth naming. Ollama measured faster on the mixture-of-experts model this time than it did in the original run, 91.23 against 83.32 at the short prompt, and even taking the older and slower figure it still beats Rapid-MLX by 14.9 percent. The 4.2x is worth taking apart, because my hardware is not the reason it does not hold up. The repository gives no machine, no model, and no method for that number, and its own body text says "up to 3x" a few paragraphs later. The benchmark it links to measures 46.7 tokens per second against Ollama's 32.1, and 58.5 against 39.0. That is 1.46x and 1.50x, on their hardware, by their method. I measured 1.40x to 1.53x on a 2022 M1 Ultra. Their work reproduces; the tagline contradicts their own data before mine enters the argument at all. I also expected Rapid-MLX to lose on prefill, because its own published benchmark volunteers that llama.cpp beats it by 1.7 to 2x on long documents. On the dense model it had the fastest time to first token of all three at both long prompts, 18.27 seconds against MLX at 18.55 and 77.21 against 77.94. On the mixture-of-experts model it was last. Prefill depends on the architecture rather than being uniformly weak, which is the same lesson as everything else on this page. ### The cache trap is not an Ollama quirk This is the part I would keep if you keep nothing else. Send an identical 16,384 token prompt twice, then send a fresh one, and watch what each engine reports for time to first token. Engine First Identical repeat Fresh prompt Ollama 90.625s 0.191s 85.433s MLX 84.652s 77.831s 78.328s Rapid-MLX 85.714s 1.159s 77.350s Rapid-MLX reports a 74x speedup for doing nothing. Ollama reports 474x. Vanilla MLX rebuilds its cache on every call and has nothing to defeat, which is exactly why a careless comparison makes it look slow against both servers. Rapid-MLX's other headline number is a 0.08 second cached time to first token, and now you can see what that measures. To be fair to them, it says cached right there in the claim, which is more than most vendor numbers manage. The point is not that anyone is cheating. It is that prompt caching is normal now, so any benchmark that sends the same prompt a few times and averages the result is timing a lookup on every engine that has one. ### The code and the data All fifty-four runs, the cache check, and the three-way harness are in a second gist. It is separate from the first one on purpose, because that gist backs the numbers already published on this page and rewriting it would strand anyone who had read them. Take the three-way code and data from this gist. And thanks to Reid for the question. It is the second time this benchmark has been improved by someone pushing back on it in public, which is the entire argument for showing your working. Here is the thread it came from. ## First, what the machine is doing Two ideas make the rest of this page make sense. If you already know them, skip to the tables. ### Reading is one job. Writing is a different one. When you send a model a prompt, it does two things in sequence, and they are not remotely alike. First it reads everything you sent. Your question, your system prompt, and the eight files your coding agent helpfully attached. This is prompt processing, or prefill. The useful thing about it is that the model already has the whole prompt, so it can chew through all of it at once. That makes it a raw arithmetic problem, and the limit is how many calculations per second the GPU can do. Then it writes a reply, one token at a time. A token is roughly three quarters of a word. This is token generation, or decode, and it has a constraint prefill does not: the model cannot guess its fourth word before it has written its third. No shortcuts. And to produce each one, it has to consult its weights, which means reading gigabytes out of memory. Every single token. So the limit on writing is not how fast the chip can calculate. It is how fast it can move data. You can do that arithmetic on a napkin. This machine moves 800GB per second, and the model I use most is about 17GB. Divide one by the other and you get roughly 47 tokens per second as a hard ceiling, before any real-world inefficiency. Nothing can beat that number, and anything claiming to is measuring something else. Why this matters here: prefill is a compute problem and decode is a memory problem. An engine can be good at one and mediocre at the other, so a single blended tokens-per-second figure tells you almost nothing. If you paste a big file and wait, that is prefill. If you sit watching words appear, that is decode. ### Quantization is like saving the model as a JPEG A model is billions of numbers, called weights. Trained at full precision, each one takes 16 bits. A 27-billion parameter model at 16 bits is about 54GB, which will not fit comfortably on most Macs and would be slow if it did, for the memory reason above. Quantization stores those same numbers in fewer bits. Eight instead of 16 halves the file. Four instead of 16 quarters it. That 27B model becomes 30GB at 8-bit and about 17GB at 4-bit, which is the difference between owning the right Mac and not. It is the same trade as saving a photo as a JPEG instead of a RAW file. Much smaller, slightly less faithful, and for most purposes you cannot tell. Models get a little less precise as you compress them, and 4-bit is where most people land because the quality cost is small and the speed gain is not. Here is the part almost nobody mentions, and it turns out to be the whole story on this page. There is more than one way to do the compressing, the two engines use different ones, and those formats do not cost the same amount of work to unpack when the model actually runs. Two files can both say 4-bit and behave differently. ## The numbers Three repetitions per cell, engines alternating so neither one eats the heat from the other. Every figure is the median. Both engines saw identical token counts. ### Qwen3.8-27B, dense, 4-bit Prompt Engine Prefill tok/s Decode tok/s First token Peak RAM 501 Ollama 225.2 21.91 2.30s 18.0GB 501 MLX 178.2 30.89 3.01s 16.4GB 3,963 Ollama 203.1 22.35 19.66s 18.2GB 3,963 MLX 215.9 29.91 18.55s 18.3GB 16,454 Ollama 193.9 18.98 85.24s 18.2GB 16,454 MLX 211.5 27.52 77.99s 20.6GB MLX generates faster at every prompt length, and it is remarkably consistent about it: 30.86, 30.92, 30.89 across three runs. Ollama wandered between 20.54 and 23.89. ### Qwen3.8-27B, dense, 8-bit Prompt Engine Prefill tok/s Decode tok/s First token Peak RAM 501 Ollama 256.9 20.41 2.02s 29.3GB 501 MLX 168.7 19.53 3.16s 29.7GB 3,963 Ollama 229.4 20.01 17.43s 29.6GB 3,963 MLX 215.5 19.20 18.57s 31.8GB 16,454 Ollama 217.8 19.06 75.86s 30.5GB 16,454 MLX 213.2 18.13 77.38s 34.0GB Same weights. Same machine. One variable changed, and a 41 percent lead became a rounding error. ### Qwen3-30B-A3B, mixture of experts, 4-bit Prompt Engine Prefill tok/s Decode tok/s First token Peak RAM 517 Ollama 1,443.4 83.32 0.39s 18.7GB 517 MLX 706.7 68.08 0.83s 17.8GB 4,175 Ollama 1,414.5 78.02 2.87s 19.0GB 4,175 MLX 1,143.2 60.16 3.57s 18.3GB 16,507 Ollama 818.6 55.85 20.41s 20.3GB 16,507 MLX 803.7 45.80 20.66s 19.5GB Ollama wins this one on both counts. But look past the comparison for a second. This model generates at 83 tokens per second where the dense 27B managed 22, in about the same memory. A mixture-of-experts model only wakes up a fraction of itself for each token, roughly 3B of its 30B parameters here, so you pay for the memory of a large model and get the speed of a small one. That is a bigger difference than anything else on this page, and it is the reason which model you pick matters more than which runtime serves it. ## Why the answer keeps moving ### The 4-bit gap is mostly not the engine Generating a token means reading every active weight out of memory. So a model with twice the bytes should generate at roughly half the speed. That is the theory, and it is usually right. Ollama decodes at 21.9 tokens per second at 4-bit, and 20.4 at 8-bit. That is 7 percent slower while reading 76 percent more data. The theory says it should have been closer to half. So Ollama's 4-bit decode is not waiting on memory. It is waiting on something else, and the obvious suspect is unpacking the weights. Q4_K_M is a mixed K-quant: the weights sit in super-blocks with their own scales and minimums, and every one has to be decoded before any arithmetic happens. MLX uses a plainer group-size-64 scheme that costs less to unwrap. There is a second tell. Ollama prefills faster at 8-bit than at 4-bit, 229 against 203 tokens per second on a 4k prompt, with a model nearly twice the size. The only way a bigger model goes faster is if the smaller one was spending its time on something other than moving bytes. MLX drops from 30.9 to 19.5 across that same change. Much closer to the halving the theory predicts, which is what you would expect from a format that is cheap to unpack. The takeaway: when someone says engine A beats engine B at 4-bit, ask which 4-bit. Two files can both say 4-bit and hold different amounts of data. ### The prefill crossover is chunk size Ollama reads your prompt in 512-token batches. That is the default in its source. mlx-lm reads it in 2,048-token steps. A 500-token prompt is one chunk either way, so all you are measuring is the overhead of getting started, and Go calling into C++ starts faster than a Python loop. Feed it 16,000 tokens and MLX makes a quarter as many trips to the GPU, each four times bigger. GPUs like that. That is the whole crossover. ## These are not the same kind of thing Half the confusion here comes from treating them as rival products. They are not. Ollama is a Go server wrapped around llama.cpp. The math happens in hand-written Metal kernels. Everything around the math is the actual product: a model registry, memory sizing you never think about, an OpenAI-compatible endpoint, and a cache that remembers your last prompt. You install it and it works. That description needs a date on it now, and the date matters for reading every Ollama number on this page. Ollama 0.19, in March 2026, put an MLX runner in the tree, and 0.33.2 ships both it and the llama.cpp one. Which runner you get is decided per model rather than per machine: Ollama picks MLX for checkpoints published in MLX format, and llama.cpp for everything else. The three checkpoints benchmarked here are GGUF k-quants, Q4_K_M and Q8_0, which the MLX runner does not implement. So every Ollama figure on this page came out of llama.cpp and Metal, on a build that also had MLX sitting right next to it. Which is worth saying plainly, because "Ollama runs on MLX now" is true and does not mean this comparison measured MLX against itself. It also sets up the obvious next experiment, and it is not one I have run: pull the same weights in MLX format, let Ollama use its own MLX runner, and see whether the mixture-of-experts result below survives. If Ollama's advantage there is a llama.cpp advantage, that test is where it shows up. MLX is Apple's array framework, closer to NumPy than to a server. It evaluates lazily, so it can see a few operations ahead and fuse them before touching the GPU, and it treats unified memory as the normal case instead of a copy to be optimized away. mlx-lm is the language model layer on top. No server. No registry. You write Python. Which explains the shape of everything above. Ollama competes on everything except the arithmetic. MLX competes on nothing but. And they are converging anyway, since the arithmetic Ollama does not compete on is now shipping inside it. This is a snapshot, not a verdict. ## The wrong answer I got first Back to that. My first run had Ollama processing a 3,932-token prompt at 19,071 tokens per second. Then 25,596. Then 30,850. Faster every time. Computers do not do that. A back-of-the-envelope check settles it. Prefill costs roughly two operations per parameter per token, so 27 billion parameters across 3,932 tokens is about 200 trillion operations. An M1 Ultra does maybe 21 trillion a second. Call it ten seconds of work. Ollama was claiming 0.14. It was reading its own cache. Send the same prompt twice and Ollama reuses the work it already did, then reports a duration covering the lookup instead of the labor. The fix is to change the first few tokens on every run: Same prompt, first run 520.5ms real work Same prompt, second run 149.7ms cache hit New prompt prefix 460.0ms real work again This is not Ollama cheating. That cache is a genuinely useful feature, and it matters more than the benchmark does, which I will get to. It just cannot be running while you measure. If you take one thing from this page, take that. It is almost certainly why you have seen Ollama prefill numbers that looked too good. Two more, if you run this yourself. Feed both engines the exact same tokens, because if each applies its own chat template you are comparing different prompts. And check the checkpoint, not the nickname: Ollama's qwen3:30b-a3b is the original release, while the MLX build most people grab is the 2507 Instruct refresh. Same name, different weights. ## So what should you actually run? Depends what you are doing with it. You sit and watch the output arrive. Switch to MLX. On a dense model at 4-bit it is worth about 40 percent more tokens per second, and that is roughly the line between reading along and waiting for the machine. Twenty tokens a second is where people start checking their phone. Thirty is where they stop noticing. You run a coding agent over a big codebase. Stay on Ollama, and it is not close. Your problem is not generation, it is the 78 to 85 seconds before the first token on a 16k prompt. An agent sends a nearly identical prompt every turn, and Ollama's cache means it pays that once instead of every time. The feature I had to switch off to measure fairly is the most valuable thing either engine does for this job. If you have not set this up yet, the Qwen and Ollama walkthrough covers the context window settings that decide whether it works at all. You need something other tools can call. Ollama. MLX has no server, so that is yours to write and maintain, to recover a difference that disappears at 8-bit anyway. You are picking a quantization. Stay at 4-bit unless you have actually measured a quality problem. Eight-bit cost me 35 percent of my generation speed and 12GB of memory, and it turns a 32GB Mac from comfortable into impossible. You want the biggest win on this page. It is not the engine. Moving from the dense 27B to the mixture-of-experts model was worth three to four times the throughput in the same memory. Pick the architecture first. Then argue about runtimes. ## Run it on your own machine The harness is public, and honestly it is the part I would want if I were reading this. It handles the traps above and reports prefill and decode separately, straight from each engine's own counters. So are the raw numbers. All fifty-four runs are in there as JSON, with the environment they ran in and every individual repetition, so you can recompute the medians in the tables above rather than take my word for them. Take the code and the data from this gist. $ uv venv --python 3.13 .venv $ uv pip install --python .venv/bin/python -r requirements.txt $ .venv/bin/python bench.py --smoke The smoke run pulls a 0.6B model and takes under a minute. One thing to check before you trust any of it: both engines should report the same prompt token count. If they do not, something is templating your prompt twice and the comparison is already dead. ## What this ran on Machine Mac Studio, M1 Ultra, 20 cores, 128GB unified memory, 800GB/s macOS 26.6.2, build 25G83 Ollama 0.33.2 MLX mlx 0.32.2, mlx-lm 0.31.3, Python 3.13.15 Method 3 runs per cell, engines alternating, medians reported, 256 tokens generated, thinking disabled The M1 Ultra is 2022 hardware, and I should be straight about what that means. These exact numbers belong to this chip. They will not transfer to an M5. What does transfer is the shape: the dequantization cost, the chunk-size crossover, and architecture mattering more than runtime. Those live in the software, not the silicon. If you are shopping for hardware rather than a runtime, the M5 Ultra buying guide covers the current generation. One limitation of every number on this page is worth naming, because I went and measured it afterwards. These are medians of three runs per cell, and three runs finish inside the first ninety seconds. On a machine with a fan that is fine. On a fanless one it is not: a MacBook Air holds only about three quarters of its cold throughput once it has been working for a few minutes, so a median of three overstates what you actually get by roughly 20 percent. I measured that separately in what a fanless Mac sustains under load. ### Join the Local AI Group Scaling localized AI workloads in enterprise and hyper-growth environments requires solving highly complex infrastructure, secure networking, and hardware optimization challenges at scale. The Local AI Group is the premier global technical network designed exclusively for active senior engineering leaders, including Chief Technology Officers, VPs of Engineering, and Directors of Engineering at Fortune 500 companies and top-tier startups. Our invitation-only space connects leaders scaling production-grade local AI systems. We bypass commercial marketing hype to focus strictly on hardware topologies, private LLM clusters, enterprise security frameworks, and custom sandboxing alongside elite peers operating at the absolute top of the global technology sector. #### Roundtable focus areas - Direct exchange on physical cluster topologies, high-throughput GPU clusters, and enterprise server architecture - Vetted blueprints for thermodynamic profiles, process orchestration, and private model deployment pipelines - Hardened boundary defense frameworks for satisfying SOC 2, ISO 27001, and GDPR perimeters with repatriated infrastructure I vet each application myself to ensure a high-signal environment of peer practitioners. Apply to Join the Slack Group Sharing confidential or proprietary information is strictly forbidden. Participation is subject to the Terms of Use. ## Frequently asked questions ### Is MLX faster than Ollama on Apple silicon? Sometimes. On a dense 27B model at 4-bit, MLX generates 41 to 45 percent faster: 30.9 tokens per second against 21.9. Switch that same model to 8-bit and the advantage disappears, with Ollama slightly ahead at 20.4 against 19.5. Switch to a 30B mixture-of-experts model and Ollama wins outright, 83.3 against 68.1. The engine matters less than the model and the quantization you picked. ### Why is MLX faster at 4-bit but not at 8-bit? Because most of the gap is the quantization format, not the engine. Ollama decodes at 21.9 tokens per second at 4-bit and 20.4 at 8-bit. That is 7 percent slower while reading 76 percent more data, which is not how a memory-bound workload behaves. The time is going somewhere else, and the candidate is unpacking the weights: Q4_K_M is a mixed K-quant with super-block scales, and MLX 4-bit is a simpler group-size-64 scheme. MLX drops from 30.9 to 19.5 across the same change, which is much closer to memory-bound. ### Which is faster for long prompts on Apple silicon? MLX, past about a thousand tokens. On a dense model Ollama leads by 26 percent at a 500-token prompt, then MLX leads by 6 to 9 percent at 4,000 and 16,000 tokens while holding roughly flat as the prompt grows. The reason is chunk size. Ollama processes the prompt in 512-token batches and mlx-lm uses 2,048-token steps, so a short prompt is one chunk either way and a long prompt lets MLX make a quarter as many trips to the GPU. ### Why do published Ollama benchmarks show impossibly fast prompt processing? Prompt caching. Ollama reuses the KV state for a repeated prompt prefix and reports a prompt evaluation duration that covers the cache hit instead of the work. Sending one 3,932-token prompt to a 27B model three times produced 19,071, then 25,596, then 30,850 tokens per second. Numbers that improve on every repetition are a cache warming up, not compute. Vary the first few tokens on every run and the effect disappears. ### Should I switch from Ollama to MLX? Switch if you run a dense model at 4-bit and you sit watching the output arrive, because 40 percent more tokens per second is the difference between reading along and waiting. Stay on Ollama if you run mixture-of-experts models, use 8-bit, or need a server other tools can call. Ollama now ships an MLX runner of its own for models published in MLX format, so the gap may close without anyone switching anything. ### How much memory does a 27B model need on a Mac? At 4-bit, 16.4GB under MLX and 18.0GB under Ollama with a short prompt, rising to about 20GB at a 16,000-token prompt as the KV cache fills. At 8-bit it is 29 to 34GB. A 32GB Mac runs 4-bit comfortably and cannot run 8-bit with a context window worth having, so 64GB is the floor if you want 8-bit. #### Building a Mac cluster for local AI 9-Part Deep Dive This article is part of an in-depth technical series detailing the creation of a localized Apple silicon server cluster for enterprise AI inference, covering Mac mini and Mac Studio hardware, local agent hosting, and agentic coding. Overview ##### How we built an M4 Mac mini cluster to cut AI cloud spend by $40k/year The business case and localized architecture that cut enterprise Google Cloud spend by $40,000 annually. Read Article Part 1 ##### Local AI use cases: local vs. cloud AI architecture The enterprise decision matrix mapping air-gapped compliance, agentic coding, robotics, batch execution, and offline operations to local Apple silicon or cloud APIs, plus the hybrid local-first framework. Read Article Part 2 ##### M5 Ultra vs. M5 Pro vs. M6 for local AI Whether to buy one 512GB M5 Ultra Mac Studio, one M5 Pro Mac mini, or a swarm of 2nm M6 Mac minis, with the memory bandwidth math that decides it. Read Article Part 3 ##### How to build an M6 or M5 Pro Mac mini cluster Step-by-step setup guide covering hardware configuration, base macOS setup, secure remote access, process management, and cloud fallbacks. Read Article Part 4 ##### Run Qwen 3.8 on Apple silicon, without rate limits Running Qwen3.8-27B locally with Ollama and Zoo Code, plus the Mac mini and Mac Studio memory bandwidth numbers that decide whether local agentic coding is usable. Read Article Part 5 ##### Best local LLMs for agentic coding on Apple silicon Qwen 3.8, Qwen3-Coder, Gemma 4, DeepSeek V4 Flash, and GLM-5.3 compared for agentic coding, with the memory math that matches each model to the Mac that runs it. Read Article Part 6 ##### Local AI agent hosting on M6 and M5 Pro Mac minis Configuring a secure, low-power private AI appliance for always-on autonomous agent workflows. Read Article Part 7 ##### Local AI Security: ISO 27001:2022, SOC 2 & GDPR Compliance Architecting a hardened physical perimeter to satisfy rigorous enterprise ISO 27001:2022 and SOC 2 audits, plus the GDPR case for keeping inference in-house. Read Article Part 8 Currently Reading ##### MLX vs Ollama on Apple silicon, measured Fifty-four benchmark runs on the same weights and matched quantization, showing where each engine wins and why the answer changes with the model. Current Page Part 9 ##### What a fanless Mac sustains under load Six ten-minute runs on an M5 MacBook Air measuring what throughput actually holds, why a median of three overstates it, and why mains power turned out slower than battery. Read Article ### Benchmarks are the easy half Knowing which runtime is faster is not the same as getting a team to run AI locally and actually trust it. Moving our own inference in-house is what led to the workshop in the first place. Every engineering leader I talked to was asking the same three questions: what does the hardware actually cost, does it survive a security review, and will the team use it. So I built a day around answering them. See AI Speaking Programs Book the Agentic Coding Workshop