Posts tagged with "AI"

Hands-On with ChatGPT Work’s New Cloud Browser Feature

Yesterday, OpenAI introduced a new ChatGPT Work feature designed to let it navigate websites that require a login without revealing your credentials to the model.

Like a lot of people, I spend far too much time clicking around websites that require a login, looking at analytics and other data, checking whether new sponsors have been booked for our podcasts, and more. It’s a tedious but necessary part of my week that slows me down and takes me away from writing and other creative work. ChatGPT’s new Work feature is designed to handle that sort of work for you.

The feature works by signing into websites using a virtual, cloud-based computer. That separates the browsing session from any browsing session on your local computer, walling the agent’s computer use off from your open tabs, cookies, browsing history, passwords, and other data. Because the login and browsing happens in the cloud, that also means work can continue whether you close the ChatGPT app or power down your computer.

According to OpenAI, an additional review model checks the sign-in request and credential destination for signs of phishing or deception. ChatGPT Work then pauses while the user signs in. Credentials entered through its secure sign-in form go directly to the cloud browser and are not visible to the model. After the user authenticates, ChatGPT resumes its work.

OpenAI also explains that its cloud-based browser doesn’t store your username and passord. Instead, it saves cookies that allow you to return to a previously authenticated session. If you don’t want to remain logged in, ChatGPT’s cloud browser settings allow you to clear browser data for individual websites or all sites.

Similar to other features in ChatGPT and Codex, users have choices when it comes to which sites ChatGPT Work can access. “Always ask” is the default, requiring the agent to check with the user before using a website, but the permission level can also be set to “Auto approve” after ChatGPT checks a site for relevancy and risk or “Always allow,” which OpenAI discourages. Individual websites can also be allowed or blocked. These settings control website access, but consequential actions require separate confirmation.

ChatGPT Work's browser control works best on a mobile device and with simple login systems.

ChatGPT Work’s browser control works best on a mobile device and with simple login systems.

I gave ChatGPT’s browser use a try and the results were mixed. I wasn’t able to get the feature to work at all using ChatGPT Work in Safari on my Mac. Some websites had security measures in place that prevented me from logging in. Other times, the cloud browser asked me to take over manually, but I was unable to do so because the UI was frozen or login panels didn’t appear.

Logging into a site with a CAPTCHA required manual intervention.

Logging into a site with a CAPTCHA required manual intervention.

I had better luck using my iPhone. On one site with a simple username and password system, a login sheet appeared. I entered my credentials, was logged in, and the agent navigated the site to answer my queries. On Apple’s affiliate link dashboard website, I had to take over the browser interaction manually to satisfy a CAPTCHA, which worked after several challenges. Once logged in, the agent pulled and analyzed the data I requested. Also, after I’d signed in to both of those sites on my iPhone, I remained logged in to ChatGPT’s cloud browser, which meant I could also use them from my Mac.

In its current form, the idea of having ChatGPT Work browse signed-in websites on your behalf is better than its implementation. If you can get logged in, having an agent collect and analyze things like analytics data is fantastic. However, getting past the initial login screen is still too frustrating. That said, I’ll be keeping a close eye on the feature, which is available on eligible plans depending on rollout and workspace settings, for future use collecting and analyzing data that would otherwise require a lot of clicking around.


The Potential of M6 and M5 Ultra for Local AI on macOS

Earlier today, Apple unveiled the new generation of Mac mini and Mac Studio, featuring the latest entries in the Apple silicon family of chips: the M6, available in the Mac mini, and the M5 Ultra, exclusive to the Mac Studio. You can read more details about the announcement and related specs in John’s overview.

As someone who’s been working with an M3 Ultra Mac Studio (on loan from Apple) with 512 GB of RAM for the past year (plus two separate M4 Mac minis) with a particular focus on local AI models and performance gains enabled by Apple’s MLX framework, I obviously am very interested in these new machines. To give you some context: my current workspace for my upcoming iOS and iPadOS 27 review, which lives in Notion, is entirely managed by a series of local agents running the latest DeepSeek-V4-Flash via MLX on the Mac Studio, which continuously scan the project for new notes, sources, and research material that is automatically categorized and linked in the chapters that I’m writing. So, yes, I’m sure I’ll have more thoughts on the new Mac minis and Mac Studios soon. In the meantime, I thought it’d be fun to break down the local AI-related details from Apple.

Read more


Defining an “Agent Harness”

I’ve recently been asked to explain what an “agent harness” is (particularly after I wrote about my favorite app of the year so far, Open Minis). I realized it was one of those concepts I could understand intuitively but not quite articulate. Thankfully, other people have done a better job of explaining it than I have.

Drew Breunig calls them “situated agents”:

Harrison Chase once excitedly shared an insight that agents are comprised of 4 things: a system prompt, a planning tool, a file system, and subagents. In the year-plus since he said that, I think this remains largely true. (Though you might tweak it to have general tools, etc.)

[…]

Imagine the simplest coding agent, and you at the keys. Let’s slowly zoom out and consider all the elements the harness can manage:

  1. The Session: The current task and context, as both a trajectory and a durable, branchable log. You can zoom backwards, fork, and replay it.
  2. The Environment: The instance, defined as a sandbox, terminal, worktree, computer, and/or container.
  3. The Repo: The project. Versioned with Git, it contains your code, history, current work, AGENTS.md, guides, and hooks.
  4. Memory: The person’s predilections, accrued over time, managing progress and past decisions.
  5. Skills: The domain, artifacts describing reusable workflows or domain knowledge worth wielding in this situation.
  6. The Team: Your colleagues and counterparts. Shared rooms, shared traces, project tracking tools, issues, and bug reports.
    • The Organization: The policies and audits, defined by legal, leadership, and procurement.
    • The Model: The LLMs themselves, the common artifact shared by all. Stochastic blobs we all poke trying to evoke positive outcomes. Log or train on their quirks and adjust.

I also thoroughly enjoyed this explanation and excellent visualization by Ted Spare, Dexter Storey, and Sarim Malik, writing for Rubric Labs:

A harness is the software that translates a model into a system that can affect its environment.

[…]

Functionally, a harness makes a model agentic, meaning it can take action.

And:

As coding agents begin to run for longer, and deploy more intelligence through dispatch, they are owning larger and more complex problems end to end, and the developer is less in the loop to steer the agent. The value of high quality planning increases as agents implement the plans more autonomously. Harnesses like Claude Code and Codex ship with a native planning mode, where the agent must first create a detailed Plan.md file with feedback from the user before executing. These harnesses then place a reference to the plan and the todo list into a top level state (system prompt) so that the agent doesn’t forget what it’s working on across long runs.

Given the multi-model, hybrid structure of the new Siri AI, we should probably assume Apple also made a lightweight “Siri harness” to aid the on-device orchestration of the entire system.

Permalink

Apple Music to Launch Labels on AI “Songs”

Ethan Millman, writing for The Hollywood Reporter:

Apple Music will soon launch labels on songs created with artificial intelligence, the company said in an email sent out to industry partners on Thursday.

In the email, obtained by The Hollywood Reporter, the streaming service said that AI music labels will come “later this year,” though Apple Music didn’t disclose a specific launch date.

The new labels come months after Apple Music launched transparency tags back in March, available for record labels and music distributors to disclose when content uploaded was “materially generated using AI.”

Good. As much as I enjoy working and tinkering with AI, AI “music” falls squarely in the category of AI implementations I do not understand and cannot even be remotely interested in. Perhaps I’m old-fashioned, but what’s the point of listening to “music” performed by something you can’t see live on a stage, or that you know has never lived a real life? (Obviously we can’t see The Beatles or Beethoven perform live today, but we know they were real humans with real motivations behind their work.)

If you ask me, AI music shouldn’t even be allowed on streaming services, but I suppose it’s too late to fix that problem by now. Hopefully Apple will also include a system-level toggle to permanently hide AI “music” from Apple Music’s UI, too.

Permalink

Hands-On with Computer History: OpenAI’s Take on Agent Memory

Late yesterday, OpenAI revealed the sort of automation catnip I love: Computer History. It’s a new feature for Pro, Business, and Enterprise subscribers that creates continuity for ChatGPT and Codex by converting the actions you take on your Mac into Markdown memory files that serve as context for subsequent interactions. The idea is that with the additional context, ChatGPT and Codex will better understand your requests, requiring you to provide less detail up front while getting better results. Anyone familiar with OpenClaw, Hermes Agent, and similar projects will know where this is going.

Here’s how it works.

Read more


The Utility App Flood Won’t Last

The App Store is evolving at a breakneck pace not seen since its earliest days. We saw the first glimmers of what was to come early in 2025, but it wasn’t until late last year that the tsunami of apps developed with the help of AI agents really took hold.

Since then, veteran developers are releasing new apps and updating existing ones faster than ever, and new developers are releasing their first apps in droves. Today, supply is dramatically up, demand is flat, and quality is seemingly simultaneously up and down, depending on where you look.

How these forces play out long-term is anyone’s guess, but it’s worth examining because, just as Federico’s link to Bryan Irace’s post about agentic coding tools foreshadowed the rise of tools like Codex and Claude Code, today’s trends are sparks that have the potential to become tomorrow’s App Store wildfires or simply fizzle out.

Utility apps are on the front lines of this change. The TechCrunch story I linked in April picked up on this trend:

Another interesting tidbit from Appfigures is that the Utilities app category moved up the top five chart.

If anything, the trend has accelerated in the months since.

It’s not surprising at all that utility apps have taken off. They’ve been a staple of new developers since long before agents came along. That’s because many are UI wrappers around command line tools. That isn’t a knock against the developers of those apps; most users don’t want to open Terminal to convert a video or audio file to another format using ffmpeg, for example. By building a great UI around command line tools, developers have made them far easier to use.

However, the relative simplicity and narrow scope of many utility apps have made them a natural fit for AI agents, too. That’s why the App Store (and my inbox) is deluged with Mac menu bar and single-screen iPhone and iPad utility apps.

On the one hand, the abundance of utilities has been great for users. More choice means you’re more likely to find the app that perfectly fits your needs.

On the other hand, though, I don’t think what’s happening in the category is sustainable and expect to see it dramatically shift again in the coming months. As we’ve seen from reporting by The New York Times, app supply is outstripping demand by orders of magnitude, which will drive down prices. That alone is likely to cause the utility app market to shrink. But there’s more to it than that.

Remember, the demand shifts that The New York Times reported, based on Sensor Tower numbers, are for the entire App Store, where download numbers have grown 2-3% this year and last. I expect utility downloads to actually shrink in the coming months – again, because of agents. Utilities, especially simple ones, are exactly the sorts of apps that are becoming trivially easy for power users – the very users these utilities are made for – to create themselves. From frontier labs‘ model improvements to a growing number of app-building tools from those same labs and third parties, it’s never been easier to build a web or native app yourself.

And although I agree with people, like Nilay Patel, who say the notion that everyone will build their own software is overblown, the utility market is different. Your average person is not downloading apps to adjust the frame rate and file format of a video or downloading YouTube videos to watch later. But those are exactly the sort of things that people who are using agents are doing with them, whether they’re having an agent do those things directly via a command line tool in an app like Codex or Claude Code or building an app to do the same thing. That’s going to put even more pressure on the category.

That said, I think there will always be a cohort of users who would rather pay for an app than build it themselves, so I’m not predicting the demise of utilities in general – just the end of today’s frothy market. I also think there remains a place for simple utilities that solve hard problems with clever solutions. Not every utility is a UI wrapper for a terminal command. Plenty of apps feature original solutions or thoughtful and unconventional remixes of disparate tools.

Utilities aren’t going away, but just like other App Store deluges, the trend will recede. It’s just that this time, it’s likely to flip faster than usual.


App Store Chaos

Kalley Huang, writing for The New York Times about apps written with the help of AI agents that are flooding the App Store:

But as with many things A.I., just because something is easy to build doesn’t mean people will use it. It is not clear if vibecoding is breathing new life into the App Store or just cluttering it.

Last year, the number of new apps released in the App Store grew 30 percent to about 600,000, according to estimates by Sensor Tower, an app analytics firm. In the first half of this year, new apps doubled to about 560,000.

Now, Sensor Tower metrics should always be taken with a grain of salt. They don’t have direct access to Apple’s sales data, so they’re extrapolating from incomplete data that they collect themselves. That said, I think it’s fair to take their numbers as broadly representative of sales trends over time, and the story they tell, as reported by Huang, is interesting.

Based on Sensor Tower’s numbers, the total number of App Store releases peaked in 2016 at 890,000, hit a low of 420,000 in 2022, and this year, is on pace to pass 2016’s peak by a healthy margin. But app releases don’t equate to downloads, let alone sales. According to Huang’s reporting, downloads increased 3% in 2025 and 2% in the first half of 2026, far lower than the growth of new releases.

All of this tracks closely with what we’ve seen at MacStories. One subtle trend I’d add is that whereas for years most developers contacted us before they released an app, a lot of new developers are doing so after their apps are on the App Store. I suspect what I’m seeing is a new generation of developers learning the hard discoverability lessons of the App Store for the first time because when I check out these apps, they rarely have any App Store reviews.

The upshot of all these statistics and trends is App Store chaos of a magnitude that we haven’t seen in a long time. It’s a little like the App Store gold rush of the early days, but without the gold. Discoverability has gone from bad to worse, and the supply of apps is off the charts compared to the demand. With download numbers barely creeping up, the flood of new apps is making selling on the App Store harder for everyone.

Yet, it’s exactly this sort of chaotic disruption that leads to exciting new apps. And, it’s tools like Codex and Claude Code that empower and democratize app development, allowing people who would never have built an app to see their ideas become a reality. Those are things I love to see.

There’s no doubt that the App Store is out of whack thanks to agent-assisted coding, but like any market it will find its equilibrium again. In the meantime, it’s never been a better time to be a fan of apps because while a lot of those record numbers are comprised of mediocre apps, there are hidden gems, and we’re on the hunt for them.

Permalink

Running Simulators Inside Claude Code for Mac

Claude Code has a new simulator button.

Claude Code has a new simulator button.

Claude Code’s Mac desktop app has added Apple simulator support, and although it’s billed as “iOS Simulator Support,” it actually works with any iOS, iPadOS, visionOS, or watchOS simulator you have installed.

The process of setting up a simulator is simple. A new button with an iPhone icon appears when you open an existing project. Click it and a new iPhone simulator pane opens to the right by default. However, clicking on the name of the simulator reveals a list of every available simulator. You can also run multiple simulators at once, but that immediately made my aging M1 Max Mac Studio start to chug a little.

With access to your running app, Claude Code is more effective at debugging UI bugs.

With access to your running app, Claude Code is more effective at debugging UI bugs.

Once you have the simulators that you want open, simply ask Claude Code to run and test your app, and it gets to work, building the app, installing it in the simulator, opening it, and poking around. I loaded up a simple word game I’ve been playing with, and after a few minutes, Claude Code found a UI bug while controlling the simulator. I asked it to fix the issue, which it did, then tested and committed. My testing has been limited, but so far it looks like a simple, effective loop.

It’s also worth noting that Claude Code is running the simulator directly inside the Claude app. This is not computer use, meaning windows aren’t going to be opening and getting in your way as you work on something else. All in all, the addition of the simulator, like the inclusion of an in-app sandboxed browser a couple of weeks ago, is a great update to Claude Code.

Permalink

Open Minis Is the iOS Agent I Wish Siri AI Could Be

Open Minis for iOS.

Open Minis for iOS.

Every once in a while, I come across a third-party app that either resets my expectations for a particular niche of software on iOS or creates an entirely new category altogether. I’ve had quite a few of these moments in the 17 years I’ve been writing app reviews at MacStories: Editorial, Workflow, Obsidian, Sky, and, most recently, OpenClaw come to mind. Over the past few weeks, I’ve had another such moment with Open Minis, a new chatbot app for iPhone and iPad that I can best describe as getting a preview of Siri AI’s agentic future, today.

Open Minis is an on-device agent that lets you use any frontier model for conversations, with a twist: unlike other AI wrappers, this app deeply integrates with all sorts of native Apple system frameworks using official APIs available to third-party developers. Open Minis can control Reminders and Music, which we have seen before, but also Calendar, Maps, HomeKit, HealthKit, and Files; it can even work with frameworks such as Vision for OCR, Apple NLP for natural language processing, the iOS clipboard, speech recognition and dictation, NFC, and Bluetooth. Furthermore, Open Minis comes with its own configurable workspace in the Files app, has an on-device memory system, and features browser-use capabilities via a (once again, native for iOS developers) built-in WebKit web view.

Effectively, Open Minis is what would happen if you rolled Claude Code and OpenClaw into one intuitive agentic experience designed for Apple users, with a powerful tool-calling harness that can improve and tweak itself, and with integrations specifically designed for the Apple ecosystem. It is the most impressive indie app I’ve seen in a while. If you were disappointed by the lack of power-user capabilities in Siri AI, this is a real, currently shipping example of what Siri AI could be if only it were powered by a frontier model with truly agentic functionalities.

Read more