This Week's Sponsor:

GoodLinks

Save articles for later and enjoy them when you have time


Posts tagged with "mac studio"

M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents

The M5 Ultra Mac Studio.

The M5 Ultra Mac Studio.

For the past few days, I’ve been testing the (currently) top-of-the-line M5 Ultra Mac Studio with 256 GB of RAM.

I’ll cut to the chase: the M5 Ultra Mac Studio is a dream machine for local AI agents. This computer makes it possible to run personal assistants powered by local models with great performance and no additional cloud costs. If you’ve been skeptical of testing OpenClaw or Hermes Agent with local models because they’d never be even remotely near the intelligence and speed of cloud ones, this Mac will change your mind about that.

Since last Thursday, I’ve been comparing this Mac Studio to its predecessor, the M3 Ultra with 512 GB of RAM, as well as my own desktop gaming PC with an RTX 5090 inside. For its size, price, thermal performance – not to mention Apple’s approach to unified memory – the M5 Ultra Mac Studio has fundamentally changed how I think about models running locally and what they can enable now. A 5090, of course, still has an edge over the M5 Ultra thanks to its higher memory bandwidth. But considering the sheer size of my PC build, as well as its heat and noise, I would prefer an M5 Ultra Mac Studio any day. It also happens to be a Mac, with an operating system that looks nice and doesn’t suck, plus a vibrant app ecosystem. (Windows fans, I’m sorry, but Microsoft software will never get my sympathy.)

As I’ll explore in this article, running the latest Qwen3.8-Flash-Next model on the M5 Ultra Mac Studio has been so nice and fast, I’ve made it my default in both Open Minis for iOS and Hermes Agent. That’s right: the personal assistants I use the most – more than Siri AI, in fact – are now entirely powered by a model running locally on a Mac Studio. Furthermore, thanks to the M5 Ultra’s faster GPU and higher memory bandwidth, these agents start responding more quickly, stay fast at larger context windows, and can run long, multi-turn loops without slowing to a crawl as the session grows. Because of this, I’ve also been using local models in the Codex app on my Mac – either as main threads or subagents orchestrated by GPT-6 Astra – and I’ve had a great experience doing so.

Local subagents running in Codex on the M5 Ultra Mac Studio.

Local subagents running in Codex on the M5 Ultra Mac Studio.

I should note upfront that I’m not an AI developer by trade: I do not train or fine-tune models. I’m a tinkerer at heart, and I’ve been playing around with local AI models for over a year at this point. This summer, I went all-in on local AI usage for a big project I was working on, which I will explain in the following section.

My goal with this article is to provide you with a mix of two things: numbers and visualizations based on the (many) tests I’ve run over the course of four days, and an explanation of my practical use cases for local AI applied to my workflow and how I get things done for MacStories.

Let’s dive in.

Read more


The Potential of M6 and M5 Ultra for Local AI on macOS

Earlier today, Apple unveiled the new generation of Mac mini and Mac Studio, featuring the latest entries in the Apple silicon family of chips: the M6, available in the Mac mini, and the M5 Ultra, exclusive to the Mac Studio. You can read more details about the announcement and related specs in John’s overview.

As someone who’s been working with an M3 Ultra Mac Studio (on loan from Apple) with 512 GB of RAM for the past year (plus two separate M4 Mac minis) with a particular focus on local AI models and performance gains enabled by Apple’s MLX framework, I obviously am very interested in these new machines. To give you some context: my current workspace for my upcoming iOS and iPadOS 27 review, which lives in Notion, is entirely managed by a series of local agents running the latest DeepSeek-V4-Flash via MLX on the Mac Studio, which continuously scan the project for new notes, sources, and research material that is automatically categorized and linked in the chapters that I’m writing. So, yes, I’m sure I’ll have more thoughts on the new Mac minis and Mac Studios soon. In the meantime, I thought it’d be fun to break down the local AI-related details from Apple.

Read more


Apple Announces New M6 and M5 Pro Mac minis, Along with M5 Max and M5 Ultra Mac Studios

Source: Apple.

Source: Apple.

Today, Apple announced a new M6 chip that will debut in the Mac mini and an M5 Ultra that is coming to the Mac Studio. The company also announced an M5 Pro Mac mini and an M5 Max Mac Studio. Let’s take a look at the headline specs.

The headline feature of the M6 chip is that it’s Apple’s first 2-nanometer chip. The chip also includes “a larger 12-core CPU complex with the world’s fastest CPU core, a larger 12-core GPU with Neural Accelerators, a Dual 16-core Neural Engine, and up to 170GB/s of unified memory bandwidth.”

Looking more closely at the M6, Apple says it consists of two super cores, four performance cores, and six efficiency cores and is 1.2x faster in multithreaded performance than an M5 and up to 2.4x faster than an M1. Apple says the GPU delivers nearly 30% more peak GPU compute for AI than the M5. The M6’s 10% faster memory bandwidth will help with AI tasks, too, but its 32GB memory ceiling will limit which models can run locally.

Source: Apple.

Source: Apple.

The M6 mini’s GPU sounds impressive, too, with two more cores, Neural Accelerators for each core, which is a first for the mini. According to Apple, that leads to an up to four-fold increase in AI performance and a 2x graphics bump compared to the M4 mini. The Neural Engine has also been improved.

Apple also introduced an M5 Pro mini. It’s understandable if you find it confusing that the base model Mac mini has a more recent generation chipset than the more expensive model. What matters more than the chip generation is the “Pro” in M5 Pro, which makes it more capable than the M6, which has been designed for entry-level models. The same sort of asymmetry occurred with the 2025 Mac Studio, which included an M4 Max model alongside the M3 Ultra. A big part of that difference is the M5 Pro mini’s support for 64GB of memory (double the M6 version’s top memory spec) and its 307GB/s of memory bandwidth compared to the M6 model’s 170GB/s memory bandwidth.

Source: Apple.

Source: Apple.

As for the M5 Ultra, which debuts in the Mac Studio, it sounds like a beast. According to Apple, the chip uses a next-generation version of its UltraFusion technology, resulting in a quad-die architecture for the first time in an Apple silicon chip. In addition, the M5 Ultra features an “up-to-36-core CPU and up-to-80-core GPU with a massive 1.2TB/s of unified memory bandwidth, 50 percent more than M3 Ultra.”

Read more


Testing DeepSeek R1-0528 on the M3 Ultra Mac Studio and Installing Local GGUF Models with Ollama on macOS

DeepSeek released an updated version of their popular R1 reasoning model (version 0528) with – according to the company – increased benchmark performance, reduced hallucinations, and native support for function calling and JSON output. Early tests from Artificial Analysis report a nice bump in performance, putting it behind OpenAI’s o3 and o4-mini-high in their Intelligence Index benchmarks. The model is available in the official DeepSeek API, and open weights have been distributed on Hugging Face. I downloaded different quantized versions of the full model on my M3 Ultra Mac Studio, and here are some notes on how it went.

Read more


Notes on Early Mac Studio AI Benchmarks with Qwen3-235B-A22B and Qwen2.5-VL-72B

I received a top-of-the-line Mac Studio (M3 Ultra, 512 GB of RAM, 8 TB of storage) on loan from Apple last week, and I thought I’d use this opportunity to revive something I’ve been mulling over for some time: more short-form blogging on MacStories in the form of brief “notes” with a dedicated Notes category on the site. Expect more of these “low-pressure”, quick posts in the future.

I’ve been sent this Mac Studio as part of my ongoing experiments with assistive AI and automation, and one of the things I plan to do over the coming weeks and months is playing around with local LLMs that tap into the power of Apple Silicon and the incredible performance headroom afforded by the M3 Ultra and this computer’s specs. I have a lot to learn when it comes to local AI (my shortcuts and experiments so far have focused on cloud models and the Shortcuts app combined with the LLM CLI), but since I had to start somewhere, I downloaded LM Studio and Ollama, installed the llm-ollama plugin, and began experimenting with open-weights models (served from Hugging Face as well as the Ollama library) both in the GGUF format and Apple’s own MLX framework.

LM Studio.

LM Studio.

I posted some of these early tests on Bluesky. I ran the massive Qwen3-235B-A22B model (a Mixture-of-Experts model with 235 billion parameters, 22 billion of which activated at once) with both GGUF and MLX using the beta version of the LM Studio app, and these were the results:

  • GGUF: 16 tokens/second, ~133 GB of RAM used
  • MLX: 24 tok/sec, ~124 GB RAM

As you can see from these first benchmarks (both based on the 4-bit quant of Qwen3-235B-A22B), the Apple Silicon-optimized version of the model resulted in better performance both for token generation and memory usage. Regardless of the version, the Mac Studio absolutely didn’t care and I could barely hear the fans going.

I also wanted to play around with the new generation of vision models (VLMs) to test modern OCR capabilities of these models. One of the tasks that has become kind of a personal AI eval for me lately is taking a long screenshot of a shortcut from the Shortcuts app (using CleanShot’s scrolling captures) and feed it either as a full-res PNG or PDF to an LLM. As I shared before, due to image compression, the vast majority of cloud LLMs either fail to accept the image as input or compresses the image so much that graphical artifacts lead to severe hallucinations in the text analysis of the image. Only o4-mini-high – thanks to its more agentic capabilities and tool-calling – was able to produce a decent output; even then, that was only possible because o4-mini-high decided to slice the image in multiple parts and iterate through each one with discrete pytesseract calls. The task took almost seven minutes to run in ChatGPT.

This morning, I installed the 72-billion parameter version of Qwen2.5-VL, gave it a full-resolution screenshot of a 40-action shortcut, and let it run with Ollama and llm-ollama. After 3.5 minutes and around 100 GB RAM usage, I got a really good, Markdown-formatted analysis of my shortcut back from the model.

To make the experience nicer, I even built a small local-scanning utility that lets me pick an image from Shortcuts and runs it through Qwen2.5-VL (72B) using the ‘Run Shell Script’ action on macOS. It worked beautifully on my first try. Amusingly, the smaller version of Qwen2.5-VL (32B) thought my photo of ergonomic mice was a “collection of seashells”. Fair enough: there’s a reason bigger models are heavier and costlier to run.

Given my struggles with OCR and document analysis with cloud-hosted models, I’m very excited about the potential of local VLMs that bypass memory constraints thanks to the M3 Ultra and provide accurate results in just a few minutes without having to upload private images or PDFs anywhere. I’ve been writing a lot about this idea of “hybrid automation” that combines traditional Mac scripting tools, Shortcuts, and LLMs to unlock workflows that just weren’t possible before; I feel like the power of this Mac Studio is going to be an amazing accelerator for that.

Next up on my list: understanding how to run MLX models with mlx-lm, investigating long-context models with dual-chunk attention support (looking at you, Qwen 2.5), and experimenting with Gemma 3. Fun times ahead!


Making a Macintosh Studio→

I’ve lost track of how many MacStories readers have sent me this over the past few days (thank you; you know me well), and, unsurprisingly, the latest project by Scott Yu-Jan is extremely my kind of thing. Scott 3D-printed a Macintosh-like shell to host a Mac Studio with an iPad mini used as its display thanks to wired Sidecar. It’s magnificent:

Obviously, as someone who relies on Sidecar on a daily basis now, I find this project a masterpiece in creativity and taking advantage of Apple’s ecosystem. I would pay serious money to have a version of this for my Mac mini and 11” iPad Pro.

Permalink

Digital Foundry Tests How a Fully-Loaded Mac Studio Stacks Up to High-End Gaming PCs→

It’s not unusual for Apple keynotes to feature gaming. Sometimes it’s about Apple Arcade, and other times it’s a demo of a third-party title coming to one of the company’s platforms. However, this year’s WWDC keynote was a little different, sprinkling developer-focused gaming announcements throughout the presentation and focusing on the upcoming release of No Man’s Sky and Resident Evil Village on the Mac. With Metal 3, controller functionality that continues to be extended, and an emphasis on titles with name recognition, many came away wondering if Apple is trying to position its latest Macs as legitimate challengers to high-end gaming PCs.

That’s the question Digital Foundry set out to answer in its latest YouTube video and companion story on Eurogamer by Oliver Mackenzie. When it comes to evaluating gaming hardware, few do it as well as Digital Foundry, which is why I was immediately curious to see what they thought of a fully loaded Mac Studio with an M1 Ultra SoC.

At just slightly larger than an Xbox Series S by volume and with ultra-low power consumption, the Mac Studio is unlike any high-performance PC. Digital Foundry came away impressed with the technical details of the M1 Ultra SoC, which held its own against high-end Intel CPUs and was in the ballpark in comparison to top GPUs:

The M1 Ultra is an extremely impressive processor. It delivers CPU and GPU performance in line with high-end PCs, packs a first-of-its-kind silicon interposer, consumes very little power, and fits into a truly tiny chassis. There’s simply nothing else like it. For users already in the Mac ecosystem, this is a great buy if you have demanding workflows.

However, the system’s performance doesn’t tell the whole story and can’t make up for the lack of videogames available for the Mac:

These results are really just for evaluating raw performance though, as the Mac is not a good gaming platform. Very few games actually end up on Mac and the ports are often low quality. If there is a future for Mac gaming it will probably be defined by “borrowing” games from other platforms, either through wrappers like Wine or through running iOS titles natively, which M1-based Macs are capable of. In the past, Macs could run games by installing Windows through Apple’s Bootcamp solution, but M1-based chips can’t boot natively into any flavour of Windows, not even Windows for ARM.

The upshot is that gaming on the Mac remains a mixed bag. Apple’s most capable M1s make the Mac more competitive with gaming PCs, but it’s not clear that the catalog of games available on the Mac will change anytime soon:

Gaming on Mac has historically been quite problematic and that remains the case right now - native ports are thin on the ground and when older titles such as No Man’s Sky and Resident Evil Village are mooted for conversion, it’s much more of a big deal than it really should be. Perhaps it’s the expense of Apple hardware, perhaps it’s the size of the addressable audience or maybe gaming isn’t a primary use-case for these machines, but there’s still the sense that outside of the mobile space (where it is dominant), gaming isn’t where it should be - Steam Deck has shown that compatibility layers can work and ultimately, perhaps that’s the route forward. Still, M1 Max and especially M1 Ultra are certainly very capable hardware and it’ll be fascinating to see how gaming evolves on the Apple platform going forward.

Digital Foundry’s results highlight that tech specs are necessary but not sufficient for videogame industry success. The Mac hasn’t been in the same league as high-end gaming PCs for a long time, and tech specs historically were just one of the issues. Given Apple’s lackluster history in desktop gaming, it’s fair to be skeptical about whether the company can attract the developers of current-generation, top-tier games to the Mac. Still, for the optimists in the crowd, the power of the M1 Ultra has brought the Mac a long way from where it stood during the Intel-baed days as a gaming platform. Personally, I’m a skeptical optimist with one foot in each camp. The hardware is heading in the right direction, but the jury’s still out on the software and Apple’s business plan to attract game developers.

Permalink

Mac Studio and Studio Display Review Roundup

The reviews are out for the Mac Studio and Studio Display and a lot has been written about both. I’ve pulled some of the most interesting tidbits from the reviews, but if you’re considering buying a Mac Studio or Studio Display, be sure to read all of these reviews because they offer a wide range of perspectives on the kind of uses for which Apple’s new hardware is best.

Read more


Mac Studio, M1 Ultra, and Apple Studio Display: The MacStories Overview

Yesterday during their Peek Performance keynote event, Apple unveiled the Mac Studio and Apple Studio Display. The former is an all-new computer joining the Mac lineup, with specs that are blowing away Apple’s previous offerings due to the introduction of a new top-of-the-line M-series chip: the M1 Ultra. The Apple Studio Display marks Apple’s true return to the consumer display market after a near decade-long hiatus.

Read more