Posts tagged with "ChatGPT"

ChatGPT Is Surprisingly Thorough at Planning Running Routes→

Every now and then, I go looking for new running routes to mix things up, which was why I was intrigued by this post by Simon Willison, who gave ChatGPT Work his address and asked it to find 5K and 10K routes, using OpenStreetMap data:

It worked for 27 minutes and produced exactly what I’d asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here’s that 5K route:

I decided to give the same thing a try and was impressed by the results. ChatGPT spent a similar amount of time, poring over OpenStreetMap data, but it didn’t stop there. The model consulted my town’s website to account for road construction and the availability of sidewalks. Then, it visited the county’s website to review maps of greenways that aren’t part of the local road system. The final results were within 100m of exact 5K and 10K routes with notes on where to be careful to verify safe pathways and road crossings.

What I love about this sort of experiment is that it’s not something I’d ever have thought to ask ChatGPT myself. In hindsight, it’s obvious that all the data is there, ready to assemble. However, it just goes to show the many ways that I, and I’m sure many others, are still just scratching the surface of what LLMs make possible.

Permalink


Hands-On with ChatGPT Work’s New Cloud Browser Feature

Yesterday, OpenAI introduced a new ChatGPT Work feature designed to let it navigate websites that require a login without revealing your credentials to the model.

Like a lot of people, I spend far too much time clicking around websites that require a login, looking at analytics and other data, checking whether new sponsors have been booked for our podcasts, and more. It’s a tedious but necessary part of my week that slows me down and takes me away from writing and other creative work. ChatGPT’s new Work feature is designed to handle that sort of work for you.

The feature works by signing into websites using a virtual, cloud-based computer. That separates the browsing session from any browsing session on your local computer, walling the agent’s computer use off from your open tabs, cookies, browsing history, passwords, and other data. Because the login and browsing happen in the cloud, that also means work can continue whether you close the ChatGPT app or power down your computer.

According to OpenAI, an additional review model checks the sign-in request and credential destination for signs of phishing or deception. ChatGPT Work then pauses while the user signs in. Credentials entered through its secure sign-in form go directly to the cloud browser and are not visible to the model. After the user authenticates, ChatGPT resumes its work.

OpenAI also explains that its cloud-based browser doesn’t store your username and password. Instead, it saves cookies that allow you to return to a previously authenticated session. If you don’t want to remain logged in, ChatGPT’s cloud browser settings allow you to clear browser data for individual websites or all sites.

Similar to other features in ChatGPT and Codex, users have choices when it comes to which sites ChatGPT Work can access. “Always ask” is the default, requiring the agent to check with the user before using a website, but the permission level can also be set to “Auto approve” after ChatGPT checks a site for relevancy and risk or “Always allow,” which OpenAI discourages. Individual websites can also be allowed or blocked. These settings control website access, but consequential actions require separate confirmation.

ChatGPT Work's browser control works best on a mobile device and with simple login systems.

ChatGPT Work’s browser control works best on a mobile device and with simple login systems.

I gave ChatGPT’s browser use a try and the results were mixed. I wasn’t able to get the feature to work at all using ChatGPT Work in Safari on my Mac. Some websites had security measures in place that prevented me from logging in. Other times, the cloud browser asked me to take over manually, but I was unable to do so because the UI was frozen or login panels didn’t appear.

Logging into a site with a CAPTCHA required manual intervention.

Logging into a site with a CAPTCHA required manual intervention.

I had better luck using my iPhone. On one site with a simple username and password system, a login sheet appeared. I entered my credentials, was logged in, and the agent navigated the site to answer my queries. On Apple’s affiliate link dashboard website, I had to take over the browser interaction manually to satisfy a CAPTCHA, which worked after several challenges. Once logged in, the agent pulled and analyzed the data I requested. Also, after I’d signed in to both of those sites on my iPhone, I remained logged in to ChatGPT’s cloud browser, which meant I could also use them from my Mac.

In its current form, the idea of having ChatGPT Work browse signed-in websites on your behalf is better than its implementation. If you can get logged in, having an agent collect and analyze things like analytics data is fantastic. However, getting past the initial login screen is still too frustrating. That said, I’ll be keeping a close eye on the feature, which is available on eligible plans depending on rollout and workspace settings, for future use collecting and analyzing data that would otherwise require a lot of clicking around.


Hands-On with Computer History: OpenAI’s Take on Agent Memory

Late yesterday, OpenAI revealed the sort of automation catnip I love: Computer History. It’s a new feature for Pro, Business, and Enterprise subscribers that creates continuity for ChatGPT and Codex by converting the actions you take on your Mac into Markdown memory files that serve as context for subsequent interactions. The idea is that with the additional context, ChatGPT and Codex will better understand your requests, requiring you to provide less detail up front while getting better results. Anyone familiar with OpenClaw, Hermes Agent, and similar projects will know where this is going.

Here’s how it works.

Read more


OpenAI Targets Coding and Knowledge Work with Its New GPT-5.5 Model

OpenAI announced GPT-5.5 and GPT-5.5 Pro today, which it says are faster and able to work more autonomously than the company’s previous models. It’s a message that is sure to interest business users whether their goal is accelerating software development or increasing productivity more generally. Some of the areas that OpenAI says GPT-5.5 and GPT-5.5 Pro excel at include:

  • writing and debugging code;
  • analyzing data;
  • conducting web research;
  • creating business documents such as spreadsheets and presentations;
  • using apps; and
  • juggling multiple tools.

In its press release, OpenAI claims that:

The gains are especially strong in agentic coding, computer use, knowledge work, and early scientific research—areas where progress depends on reasoning across context and taking action over time. GPT‑5.5 delivers this step up in intelligence without compromising on speed: larger, more capable models are often slower to serve, but GPT‑5.5 matches GPT‑5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence. It also uses significantly fewer tokens to complete the same Codex tasks, making it more efficient as well as more capable.

I haven’t tried either model yet, but early reactions seem to support OpenAI’s claims that GPT-5.5 understands user intent better, requiring less precise instructions. The company says it is better at using the tools at its disposal, and checking its own work, too. OpenAI says the Pro model takes that up a notch, working faster on more complex tasks, such as programming, research, and document-intensive workflows. Whether the early hype translates into real-world gains that are noticeable in everday work, remains to be seen, but we shouldn’t have long to wait though, since GPT-5.5 is rolling out to users now.

GPT-5.5 is available in ChatGPT and Codex to Plus, Pro, Business, and Enterprise subscribers, and GPT-5.5 Pro is limited to Pro, Business, and Enterprise subscribers in ChatGPT. Neither model is available through OpenAI’s API, but the company says they will be soon.


Roadtripping with ChatGPT Voice Mode

On Saturday, my wife Jennifer and I drove to Blowing Rock, a quaint little town in the Blue Ridge Mountains. We’d been there once before, but didn’t know the town well, so as we headed west I poked at the ChatGPT icon on my dashboard to give the app’s new CarPlay integration a try. I asked:

What activities would you recommend for a day trip to Blowing Rock, North Carolina?

What I got back was a short but good list of highlights including a hike, a visit to the Blowing Rock cliffside overlook, a few restaurants, a coffee shop, and some local shops. It was similar to a list of activities I’d looked up before we left using Claude. So far, so good.

I switched back to Apple Maps and was thinking I probably wouldn’t use ChatGPT in my car very often, but that it could come in handy for similar requests, when things got a little creepy. I explained to Jennifer that ChatGPT’s CarPlay feature was new, and I had been meaning to check it out all week. Then, just as I’d said I thought it had done a pretty good job, a voice interrupted. It was ChatGPT’s voice mode saying it was glad I liked it.

You see, just like a phone call doesn’t drop when you switch apps in CarPlay, neither does ChatGPT. I supposed I should have anticipated that the mic would remain live, but I didn’t. Nor did I notice the End button in the corner of the screen; I was driving, not studying the app’s UI.

I take it as a positive sign that I didn’t expect ChatGPT to follow me back to Apple Maps. I treat chatbots like I do any app. Give it some input, and you get an output. Close the app, and you’re done. It’s not my little robot buddy. It’s a tool like any other app.

Of course, that’s not how the voice modes of these chatbots are designed to work. Chats are meant to be an engaging back and forth. But having ChatGPT jump in on our one-on-one conversation while driving down the highway was too much. Suddenly, it felt like something else was in the car eavesdropping on us.

The experience was a good lesson in the balancing of utility and social norms around AI tools. Useful as they can be in some situations, their developers need to be more mindful of user expectations and provide better cues about how they work to avoid uncomfortable surprises. The recommendations we got from ChatGPT were good, but I also don’t expect it will get a second chance on our family road trips anytime soon.


Adobe Announces Image and PDF Integration with ChatGPT

Source: Adobe.

Source: Adobe.

Adobe announced today that it has teamed up with OpenAI to give ChatGPT users access to Photoshop, Express, and Acrobat from inside the chatbot. The new integration is available starting today at no additional cost to ChatGPT users.

Source: Adobe.

Source: Adobe.

In a press release to Business Wire, Adobe explains that its three apps can be used by ChatGPT users to:

  • Easily edit and uplevel images with Adobe Photoshop: Adjust a specific part of an image, fine tune image settings like brightness, contrast and exposure, and apply creative effects like Glitch and Glow – all while preserving the quality of the image.
  • Create and personalize designs with Adobe Express: Browse Adobe Express’ extensive library of professional designs to find the best one for any moment, fill in the text, replace images, animate designs and iterate on edits – all directly inside the chat and without needing to switch to another app – to create standout content for any occasion.
  • Transform and organize documents with Adobe Acrobat: Edit PDFs directly in the chat, extract text or tables, organize and merge multiple files, compress files and convert them to PDF while keeping formatting and quality intact. Acrobat for ChatGPT also enables people to easily redact sensitive details.
Source: Adobe.

Source: Adobe.

This strikes me as a savvy move by Adobe. Allowing users to request image and PDF edits and design documents with natural language prompts makes its tools more approachable. That could attract new users who later move to an Adobe subscription to get more control over their creations and Adobe’s other offerings.

From OpenAI’s standpoint, this is clearly a response to the consumer-facing Gemini features that Google has begun releasing, which include new image and video generation tools and reportedly caused Sam Altman to declare a “code red” inside the company. I understand the OpenAI freakout. Google has a huge user base and has been doing consumer products far longer than OpenAI, but I can’t say I’ve been very impressed with Gemini 3. Perhaps that’s simply because I don’t care for generative images and video, but these latest moves by Google and OpenAI make it clear that they see them as foundational to consumer-facing AI tools.


Why is ChatGPT for Mac So Good?→

Great post by Allen Pike on the importance of a great app experience for modern LLMs, which I recently wrote about. He opens with this line, which is a new axiom I’m going to reuse extensively:

A model is only as useful as its applications.

And on ChatGPT for Mac specifically:

The app does a good job of following the platform conventions on Mac. That means buttons, text fields, and menus behave as they do in other Mac apps. While ChatGPT is imperfect on both Mac and web, both platforms have the finish you would expect from a daily-use tool.

[…]

It’s easier to get a polished app with native APIs, but at a certain scale separate apps make it hard to rapidly iterate a complex enterprise product while keeping it in sync on each platform, while also meeting your service and customer obligations. So for a consumer-facing app like ChatGPT or the no-modifier Copilot, it’s easier to go native. For companies that are, at their core, selling to enterprises, you get Electron apps.

I don’t hate Electron as much as others in our community, but I can’t deny that ChatGPT is one of the nicest AI apps for Mac I’ve used. The other is the recently updated BoltAI. And they’re both native Mac apps.

Permalink

Apps in ChatGPT→

OpenAI announced a lot of developer-related features at yesterday’s DevDay event, and as you can imagine, the most interesting one for me is the introduction of apps in ChatGPT. From the OpenAI blog:

Today we’re introducing a new generation of apps you can chat with, right inside ChatGPT. Developers can start building them today with the new Apps SDK, available in preview.

Apps in ChatGPT fit naturally into conversation. You can discover them when ChatGPT suggests one at the right time, or by calling them by name. Apps respond to natural language and include interactive interfaces you can use right in the chat.

And:

Developers can start building and testing apps today with the new Apps SDK preview, which we’re releasing as an open standard built on the Model Context Protocol⁠ (MCP). To start building, visit our documentation for guidelines and example apps, and then test your apps using Developer Mode in ChatGPT.

Also:

Later this year, we’ll launch apps to ChatGPT Business, Enterprise and Edu. We’ll also open submissions so developers can publish their apps in ChatGPT, and launch a dedicated directory where users can browse and search for them. Apps that meet the standards provided in our developer guidelines will be eligible to be listed, and those that meet higher design and functionality standards may be featured more prominently—both in the directory and in conversations.

Looks like we got the timing right with this week’s episode of AppStories about demystifying MCP and what it means to connect apps to LLMs. In the episode, I expressed my optimism for the potential of MCP and the idea of augmenting your favorite apps with the capabilities of LLMs. However, I also lamented how fragmented the MCP ecosystem is and how confusing it can be for users to wrap their heads around MCP “servers” and other obscure, developer-adjacent terminology.

In classic OpenAI fashion, their announcement of apps in ChatGPT aims to (almost) completely abstract the complexity of MCP from users. In one announcement, OpenAI addressed my two top complaints about MCP that I shared on AppStories: they revealed their own upcoming ecosystem of apps, and they’re going to make it simple to use.

Does that ring a bell? It’s impossible to tell right now if OpenAI’s bet to become a platform will be successful, but early signs are encouraging, and the company has the leverage of 800 million active users to convince third-party developers to jump on board. Just this morning, I asked ChatGPT to put together a custom Spotify playlist with bands that had a similar vibe to Moving Mountains in their Pneuma era, and after thinking for a few minutes, it worked. I did it from the ChatGPT web app and didn’t have to involve the App Store at all.

If I were Apple, I’d start growing increasingly concerned at the prospect of another company controlling the interactions between users and their favorite apps. As I argued on AppStories, my hope is that the rumored MCP framework allegedly being worked on by Apple is exactly that – a bridge (powered by App Intents) between App Store apps and LLMs that can serve as a stopgap until Apple gets their LLM act together. But that’s a story for another time.

Permalink