> ## Content Index
> Fetch the complete content index at: https://www.siliconsnark.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Deep Dive into OpenAI DevDay 2026 Rumors
- URL: https://www.siliconsnark.com/openai-devday-2026-rumors-o-agent-pro-max-ultrafast/
- Published: 2026-09-27T01:28:59.000Z
- Updated: 2026-09-27T01:28:59.000Z
- Description: OpenAI DevDay 2026 could bring an always-on agent, faster AI, and a $500 Pro Max plan. We sort the September 29 rumors from the subscription fan fiction.
- Author: CircuitSmith
- Tags: AI, OpenAI, Deep Dive, DevDay, Guides

An always-on agent called “o,” a rumored $500 Pro Max plan, very fast tokens, and the eternal promise of One More Model. Here is what September 29 might actually deliver, what is already shipping, and which predictions need to be escorted out of the group chat.

The most believable thing about the OpenAI DevDay rumor mill is that the future might cost $500 a month.

Not because that price is confirmed. It is not. But because an industry that promised to make intelligence abundant has developed an extraordinary talent for making the good button a separate subscription. We went looking for a universal assistant and somehow wound up at the airline counter, learning that reasoning is included but overhead-bin reasoning requires Pro Max.

On Tuesday, September 29, OpenAI will hold DevDay 2026 at Fort Mason in San Francisco. The [official event page](https://devday.openai.com/?ref=siliconsnark.com) lists an opening keynote featuring Sam Altman at 10 a.m. Pacific, or 1 p.m. Eastern, with a free livestream. In-person applications are closed; registration for invited attendees is $650\. Technical sessions, demos, workshops, and a closing session fill out the day. Other sessions will be posted afterward.

Those are facts. Beyond them lies an unusually busy departure lounge for speculation.

The substantive rumors involve a persistent assistant reportedly called “o,” a higher-priced ChatGPT subscription, wider access to very fast inference, and changes that could turn OpenAI’s developer platform into a more complete place to build software. Then there are the ceremonial offerings: a bigger secret model, an AI device, video generation, and the suggestion that the next stage presentation will finally settle the nature of intelligence before lunch.

My read: the most consequential potential announcement is a better arrangement of the parts OpenAI already has. Intelligence, tools, a place to run, a way to remember the job, and a meter. Put them together well and you get useful delegated work. Put them together badly and you get an expensive intern who has discovered recursion.

*This preview reflects sources checked September 26, 2026\. Rumors remain rumors; forecasts below are editorial judgments, not inside information.*

## First, Check Whether the “Leak” Already Has Release Notes

A good DevDay preview must begin by preventing September from announcing itself twice.

OpenAI’s [September release notes](https://help.openai.com/en/articles/6825453-chatgpt-release-notes?ref=siliconsnark.com) already record GPT-6 Astra’s introduction on September 3, GPT-6 Sol and Luna arriving in Work and Codex on September 22, and plugins in Voice plus Voice in Work on September 23\. The Sol and Luna entry explicitly distinguishes those models from the ones available in ordinary Chat. The same running changelog records a planned retirement of custom GPTs and migration toward plugins.

That is a lot of product movement before anyone has put on a conference lanyard. A prediction that OpenAI will “introduce GPT-6” on Tuesday needs considerably more specificity. Introduce it where? In which product? Under what access conditions? With which model variant? The answer cannot simply be “yes, but with louder lighting.”

These distinctions sound pedantic until you are the person paying for a feature you assumed was included. An AI product now has a model, a mode, a surface, a plan, and possibly a speed class. Ordering lunch in an airport is less complicated, and that involves distinguishing three separate businesses called Express.

We covered the Astra moment in [our September launch piece](https://www.siliconsnark.com/openai-launches-gpt-6-astra-declares-agi-and-asks-it-not-to-touch-anything/). Tuesday should be judged on what changes from that baseline.

The useful question is whether more people can accomplish more complete work with less friction. Another model name might help. So might an unglamorous improvement that stops a task from forgetting why it opened seventeen browser tabs. One earns a standing ovation; the other earns repeat customers.

## The Scorecard, Before Everyone Starts Seeing Prophecy in a Loading Spinner

Here is the map for the rest of this expedition. These labels measure the evidence behind a prediction, not a mathematically calculated probability that Sam Altman will say a particular noun.

| Watch item                                  | Evidence today                                                                | My DevDay expectation                                       |
| ------------------------------------------- | ----------------------------------------------------------------------------- | ----------------------------------------------------------- |
| Persistent agents and orchestration         | Official infrastructure already exists; reported unreleased product clues     | Strong thematic expectation; exact launch package uncertain |
| “o” / Aeon                                  | Reported configuration and upgrade-page references, plus social-source claims | Credible candidate, unconfirmed name, access, and timing    |
| Wider Ultrafast access                      | Official limited preview; reported hidden platform controls                   | One of the better-supported rollout possibilities           |
| Pro Max at $500 monthly                     | Reported unreleased subscription references                                   | Watch closely; no confirmed commercial offer                |
| Free, Prototype, Accelerate developer plans | Reported onboarding references                                                | Plausible; full application hosting remains an inference    |
| Bel, Astra 6.1, another giant model         | Social speculation without a verified DevDay release commitment               | Possible surprise, weak foundation for a confident forecast |
| Hardware or video showcase                  | Broader ambitions and circulating predictions                                 | Wild cards, especially on availability dates                |

There is a reason the table contains so many qualifying words. A configuration string can show that someone built or tested something. It cannot prove that the product survived launch review, that its name survived marketing, or that its price survived a spreadsheet populated by people with competing theories of abundance.

Multiple stories can also share one underlying source. Three articles repeating one social post are not three confirmations. They are one rumor wearing a conference badge, a press badge, and a false mustache.

## “o” Might Be Coming. The Alphabet Has Retained Counsel.

The most interesting product rumor is the smallest possible branding exercise.

In its September 26 [report on an always-on agent](https://www.testingcatalog.com/openai-to-announce-o-always-on-agent-during-devday/?ref=siliconsnark.com), TestingCatalog says it found “O” as a display name in ChatGPT configuration, an “-o” email suffix, and a reference to the agent on an upgrade page for the $100 Pro plan. It connects those clues to possible persistent-assistant functionality and the name Aeon. That is reported evidence of development, not an official announcement or an entitlement guarantee.

Separately, [RuntimeWire traced an earlier claim](https://runtimewire.com/article/openai-s-rumored-o-agent-points-to-a-long-running-grok-bot-like-product?ref=siliconsnark.com) to the X account @synthwavedd: an agent called “o,” allegedly powered by an Astra variant called “aeon,” aimed at long-running tasks and potentially DevDay. RuntimeWire labels it unconfirmed. The relationship between those names, the eventual underlying model, and any public product is unresolved.

As branding, “o” is magnificently economical. It could be a revolutionary assistant, a typo, a surprised expression, or the noise you make after reading the invoice. Search engines will love being asked to distinguish it from the fifteenth letter.

But the product concept matters. A persistent assistant could move the interaction from “answer this while I watch” toward “keep progressing on this and return when there is something worth my attention.” The appeal is easy to understand. Nobody dreams of spending their career re-explaining a project to a rectangle.

My expectation is stronger for the category than for the name. Something aimed at continuing work would fit the public direction. Whether Tuesday delivers a consumer product, a preview, a business feature, or an elegantly typeset waitlist is still open.

## Always-On Is Easy. Knowing When to Shut Up Is the Premium Feature.

Consider a hypothetical assignment: watch three suppliers, flag meaningful price changes, keep a comparison current, and prepare a recommendation before Friday. Do not buy anything.

A useful persistent assistant needs to remember the suppliers, distinguish a real price change from a promotional banner, preserve the latest comparison, notice a dead page, and respect the boundary around purchases. It also needs to understand that “keep me posted” does not mean “interrupt me twelve times to announce that Thursday still exists.”

That is the practical bar for this category, not a claim about the unreleased “o.” The demo will probably focus on initiative. I would focus on judgment.

Can it tell a failed attempt from a completed task? Can you change the objective halfway through? Can it explain what it did while you were away? Can you revoke access without wandering through a hedge maze of settings named after abstract virtues?

Our [guide to computer-use agents](https://www.siliconsnark.com/computer-use-agents-explained-why-openai-anthropic-and-perplexity-want-to-operate-your-laptop/) is useful background here. Once software can interact with other software on your behalf, the interface stops being merely a place to type. It becomes a place to delegate authority.

I want the boring demonstration: interrupt the agent. Change a requirement. Close the laptop. Reopen the task somewhere else. Have one connected service fail. Watch what survives.

The perfect happy-path demo is the wedding photo. The interruption demo is finding out whether the marriage can survive assembling furniture. For an assistant meant to stick around, the second one tells you considerably more.

## The Plumbing Is Already Here, Wearing a Public Beta Badge

There is a sturdy reason to expect agents to dominate the event. OpenAI [introduced the Agents API on September 10](https://openai.com/index/introducing-the-agents-api/?ref=siliconsnark.com) in public beta. The announcement describes an OpenAI-managed Codex harness and infrastructure for long-running agents, including context management, tools, and subagent coordination. Developers can choose execution environments, including OpenAI-hosted sandboxes and other infrastructure.

A harness is the surrounding machinery that lets a model do a job: decide what to try, use a tool, examine the result, keep useful context, and continue. Calling it a harness is unusually honest terminology for an industry otherwise inclined to call a loop “an autonomous cognitive workforce.”

The model supplies capability. The rest of the system determines whether that capability can survive contact with your actual files.

A concrete example: asking for a sales analysis can involve finding a spreadsheet, checking column definitions, cleaning dates, making a chart, discovering that one region uses a different currency, and revising the conclusion. A beautiful paragraph produced at step two is not the completed assignment. It is a confident interruption.

My inference is that DevDay could package this existing foundation into experiences that require less assembly. That could mean better setup, management, distribution, or application-building tools. None requires a miraculous new model to be meaningful.

Readers of our [AI coding agents deep dive](https://www.siliconsnark.com/deepdai-coding-agents-explained-why-every-software-company-now-wants-a-robot-engineer-on-payroll/) will recognize the stakes: the product is increasingly the environment around the model. The engine matters. So does having a car instead of four wheels laid out around an inspirational keynote.

## Ona Is a Better Clue Than Somebody Posting Three Fire Emojis

In June, OpenAI [announced an agreement to acquire Ona](https://openai.com/index/openai-to-acquire-ona/?ref=siliconsnark.com), explicitly tying the deal to persistent, customer-controlled cloud environments for Codex and work lasting hours or days. The announcement described security boundaries, scoped credentials, logging, and review. It also said the transaction remained subject to customary closing conditions. That announcement alone does not establish that the acquisition has closed.

It also does not prove Ona is secretly “o.” Similar letters are not due diligence. Otherwise Oracle would own the moon.

What it does establish is a public strategic interest in work that continues beyond a single device or active session. That is a more substantial signal than decoding a product executive’s punctuation.

The commercial logic is straightforward. If a person must keep their laptop awake to supervise every meaningful action, the agent remains attached to the person’s workday. If execution can continue in a durable environment with appropriate boundaries, the service becomes something a team can build a recurring process around.

Imagine an overnight maintenance job that leaves a reviewable report and proposed changes. The value comes from getting a trustworthy result the next morning. Whether a robot avatar blinked thoughtfully for six hours is largely an artistic matter.

The more ambitious the delegated work, the more important its housing becomes. Where does it run? Who controls its keys? What survives a restart? Who can stop it?

These are not glamorous questions. They are the questions that determine whether a company buys a system or lets one enthusiastic employee expense it until somebody in IT notices.

## Ultrafast Is Real. The Rumor Is Who Gets It Next.

This is where the preview can leave the fog machine and touch a documented product.

OpenAI’s [August 13 Ultrafast announcement](https://openai.com/index/previewing-ultrafast/?ref=siliconsnark.com) described a limited API preview running GPT-5.6 Sol on Cerebras infrastructure at up to 750 output tokens per second, and up to 14 times the speed of Standard processing. Those are OpenAI’s advertised upper-bound figures for that model and service tier. They are not a guarantee that every task, every model, or every account suddenly operates fourteen times faster.

The new clue is TestingCatalog’s September 26 [report of a hidden Playground speed selector](https://www.testingcatalog.com/openai-prepares-to-expand-ultrafast-api-to-more-users/?ref=siliconsnark.com) with Standard, Fast, and Ultrafast options. It makes broader access a plausible DevDay announcement. It does not establish general availability or universal support across GPT-6 models.

Speed sounds like a cosmetic upgrade if you think of AI as a paragraph vending machine. An answer appears more quickly; delightful, now we can disagree with it before the kettle boils.

For an agent taking many sequential steps, the effect can be more consequential. Waiting repeats. A few seconds saved on each decision can add up across a workflow. Shorter loops also make it easier for a human to stay engaged, correct the direction, and run another attempt.

My strongest expectation here is more detail about access, supported models, and economics. The impressive number is already in the public record. The useful announcement would explain how ordinary builders can purchase the thing without applying to a small society of Extremely Fast People.

## Your Browser Still Has to Load the Website, Tragically

Here is the part of the speed story the confetti cannon tends to obscure: generating tokens is only one component of doing work.

Suppose a hypothetical task takes ten minutes. Two minutes are model generation; eight are downloads, page loads, code execution, and other waiting. Make generation ten times faster and the overall job still takes eight minutes and twelve seconds. That is an improvement, but it is not a time machine. The website has declined your invitation to transcend physics.

This matters when watching an Ultrafast demo. Ask where the time went before the upgrade. A workflow dominated by model output can benefit dramatically. One dominated by a sluggish third-party system may benefit much less. A task waiting on human approval can spend all afternoon performing the most advanced operation in enterprise computing: waiting for Gary.

The other distinction is speed versus correctness. Faster generation does not, by itself, establish better judgment. If an agent takes the wrong path, accelerating it may simply produce the wrong deliverable before you have formed a coherent objection.

There is still a good product opportunity. A system that lets users choose speed according to the job could be useful. Interactive debugging and a report due next week do not necessarily need identical economics.

What I would like on Tuesday is a completed-work comparison: same task, same quality bar, actual elapsed time, actual cost. Show the awkward pauses. Include the retries.

A tokens-per-second chart is a speedometer. I would also like to know whether the vehicle arrived at the correct building.

## Pro Max, or How to Give a Subscription a Vehicle Trim Level

The money rumor comes from TestingCatalog’s September 24 [report of unreleased Pro Max references](https://www.testingcatalog.com/openai-prepares-new-500-month-pro-max-plan-for-chatgpt/?ref=siliconsnark.com). It describes a proposed $500 monthly price and the phrase “Fastest Work and Codex.” The report says the $600 shown in its video includes VAT. Higher allowances and a direct Cerebras connection remain unconfirmed, as does the actual launch offer.

That distinction matters. A product string is evidence to examine, not an invoice OpenAI has sent you.

Still: Pro Max. Of course. “Pro” apparently no longer conveys sufficient professionalism. Now one must be maximally professional. Eventually the hierarchy will reach Pro Max Executive Reserve, with a tiny artificial sommelier recommending which reasoning effort pairs best with your quarterly forecast.

If the reported price survives, the annual arithmetic is $6,000 before applicable taxes. For an individual asking occasional questions, that would be a startling amount of money. For a professional whose bottleneck is expensive, time-sensitive work, it might be defensible. Price alone does not settle value.

The package does. Which models? Which speeds? What allowance? What happens after that allowance is exhausted? Is there a usable estimate before an ambitious assignment begins?

These are the questions any eventual offer must answer. A word like “fastest” can describe an excellent service and still leave almost everything a buyer needs to know unstated.

My forecast: pricing and access deserve as much attention as the keynote’s cleverest demo. The next era of AI may arrive with a breakthrough. It may also arrive with the revelation that your existing plan was merely emotionally professional.

## The Only Benchmark Your Accountant Will Recognize

The reasonable way to assess an expensive AI plan is cost per useful completed job, including the human time required to check and repair it.

Here is deliberately simple hypothetical arithmetic. A $500 monthly tool that saves ten valuable working hours has a subscription cost of $50 per hour saved. If it saves twenty, that becomes $25\. If it produces beautiful output that takes as long to fix as doing the assignment yourself, the hourly savings are imaginary and the subscription remains extremely real.

Those examples are not measured returns for Pro Max, which has not been confirmed. They are a way to avoid buying a comparative adjective.

You must also distinguish time recovered from revenue created. Finishing a proposal faster does not automatically create another client. Saving an engineer an hour does not always remove an hour from payroll. Sometimes it creates room for better work; sometimes it creates room for another meeting, because organizations are skilled at converting efficiency into calendar debris.

Our investigation into [whether AI agents actually make money](https://www.siliconsnark.com/do-ai-agents-actually-make-money-in-2026-or-is-it-just-mac-minis-and-vibes/) is the relevant companion reading. An exciting capability and a repeatable economic outcome are different milestones.

I would be happy to see OpenAI show a task ledger: the assignment, the total spend, the result, and the amount of human intervention. No faux employee salary comparison. No assumption that every generated word represents reclaimed billable labor.

Just evidence that a specific job became cheaper or easier. Revolutionary concept. We could call it Accounting, then announce Accounting Ultra when it supports more than one column.

## The Developer Platform May Be Applying to Become Your Landlord

One of the quieter rumors could have the largest implications for builders.

TestingCatalog’s September 23 [report on new platform onboarding](https://www.testingcatalog.com/devday-new-plans-and-new-platform-for-building-ai-apps/?ref=siliconsnark.com) describes three unreleased labels: Free, Prototype, and Accelerate. The reported flow positions them around experimentation, small applications, and production workloads, with Accelerate starting at $50\. Crucially, the article says the discovered material does not confirm complete application deployment or hosting. There is no visible rollout date in that report.

So “OpenAI launches its own app-hosting platform Tuesday” would outrun the evidence. “OpenAI may be simplifying the path from API experiment to production” is a more defensible read.

The strategic attraction is obvious. A developer who begins with model access might also need execution environments, secrets, storage, usage controls, and a way to ship the result. Every boundary is a chance for the project to get stuck or for another company to win the relationship.

A well-integrated platform can make that journey easier. It can also make departure more complicated. Convenience is wonderful until you discover that six different components of your product now share the same landlord.

Tuesday’s important distinction would be between a tidier billing screen and a materially broader development environment. Both can be useful. They are not the same announcement.

I would watch for working deployment examples, export paths, and clarity about which components run where. “Build anything” is a lovely promise. “Here is how you move your database if you change your mind” is an even lovelier sentence, although it rarely gets projected in forty-foot letters.

## DevDay Has Always Been a Real-Estate Presentation for Developers

The historical pattern helps explain why the platform rumors deserve attention.

At [DevDay in November 2023](https://openai.com/index/new-models-and-developer-products-announced-at-devday/?ref=siliconsnark.com), OpenAI announced GPT-4 Turbo with a 128K context window, lower prices, the Assistants API, and additional multimodal capabilities. OpenAI’s [account of DevDay 2025](https://developers.openai.com/blog/codex-at-devday?ref=siliconsnark.com) describes launches including the Apps SDK and AgentKit. The company has repeatedly used the event to expand what developers can build around its models.

My interpretation is that the event is partly a negotiation over where future software will live. A model release offers capability. A platform release proposes an address.

The pitch is attractive: bring your idea, use our tools, reach users, and let us absorb some of the annoying infrastructure work. For a small team, removing that overhead can change what is possible. It deserves more respect than a reflexive complaint that every platform is a trap.

But a developer conference is also a sales event. The host wants your future roadmap to depend on its present roadmap. Nobody spends this much on stage design to announce that you should keep all your options equally open.

The practical response is neither worship nor cynicism. Ask what you gain, what you surrender, and what happens when the names change again.

A good DevDay would make the development path easier to understand. A less good one would add a new layer of abstraction and require developers to attend next year’s conference to learn which part was deprecated.

The most useful souvenir is working documentation. The tote bag is merely where you put your unresolved migration questions.

## Codex Needs to Show the Part After “Ta-Da”

I have already written [a very affectionate column about Codex](https://www.siliconsnark.com/openai-codex-deserves-flowers-preferably-delivered-by-a-passing-build-agent/). This is not an intervention from someone who believes all generated code should be returned to the sea.

My interest in Tuesday’s coding demonstrations is what happens after the first impressive result. A generated interface can look excellent. A test can pass. A feature can still misunderstand the request in a way that becomes obvious only when a real person uses it.

A strong demonstration would show the system inspecting existing work, making a bounded change, testing it, explaining uncertainty, and leaving a result that a person can evaluate efficiently. Better still: show it encountering a misleading test and correcting course.

OpenAI’s [September customer story about Cognition and Devin](https://openai.com/index/cognition-devin-testing-with-astra/?ref=siliconsnark.com) emphasizes Astra’s ability to test work and show results. That is a vendor-hosted customer account, not independent proof of universal reliability, but it identifies exactly the right area to scrutinize.

Code generation is only useful inside a larger chain of decisions. Someone must decide whether the change solves the right problem, whether it breaks something else, and whether it belongs in the product.

The dream is that stronger agents make those checks easier. The nightmare is that they increase output so quickly that review becomes an all-you-can-eat buffet staffed by one exhausted person with a fork.

If Tuesday includes better review, interruption, or coordination features, do not dismiss them because they lack a new model name. Removing a bottleneck is a legitimate breakthrough. It simply photographs less well than a benchmark with a bar extending confidently off the slide.

## Bel, Astra 6.1, and the Sacred Tradition of the Secret Bigger One

No AI conference is complete without a rumored model behind the rumored model, waiting in a secure facility to justify somebody’s entire online personality.

The “Bel” speculation has circulated in [a Reddit post reproducing claims attributed to @synthwavedd](https://www.reddit.com/r/ChatGPT/comments/1vztv6r/openais%5Fnext%5Fpretrain%5Fcodename%5Fbel%5Fdevday/?ref=siliconsnark.com), describing a large successor pretraining effort. The thread’s DevDay connection is speculative. I could not verify a public OpenAI commitment to announce Bel at this event. Nor did the sources reviewed establish an Astra 6.1 launch date.

The useful distinction is between an internal research project and a releasable product. A training run can exist without having a public name, stable serving economics, completed evaluations, acceptable reliability, or a support page that explains why your account cannot see it yet.

Parameter counts, when unverified, add scientific-looking decoration without solving any of those problems. An enormous number is not a release schedule. It is sometimes just a very large number wearing a lab coat.

There could absolutely be a model surprise. The point is to keep the expectation proportional to the evidence. A forecast should not turn “someone says a thing is being developed” into “Tuesday, 10:17 a.m., humanity changes forever, refreshments immediately afterward.”

For orientation amid the naming turbulence, keep our [guide to GPTs, Claude, Gemini, and friends](https://www.siliconsnark.com/the-definitive-guide-to-ai-gpts-and-friends/) nearby. Then judge any new model on availability, task performance, and cost.

My stance: leave room for a surprise. Do not let an unverified codename become the minimum acceptable outcome of a developer conference. That is how useful products get declared failures for declining to become a religion.

## Anthropic Has Entered the Chat, Carrying a Calculator

There is real competitive context behind the model anxiety.

Anthropic [introduced Claude Opus 5.5 on September 22](https://www.anthropic.com/claude-opus-5-5?ref=siliconsnark.com). Its announcement claims improved performance and estimates typical token-billed workloads cost 40 percent less than Opus 5\. It lists $4 per million input tokens and $20 per million output tokens, alongside cheaper cache reads. Those are Anthropic’s pricing and performance claims, not a head-to-head verdict from this article.

The significant point is that the contest involves economics as well as intelligence. A rival does not need to win every benchmark to win a workload. It can be more practical, more predictable, faster at the relevant task, or easier to fit into an existing process.

That makes Tuesday’s possible combination of agents, speed, and platform tooling more interesting than a single leaderboard duel.

It also makes internet explanations of release timing less trustworthy. Two companies shipping in the same week does not establish that one forced the other’s hand. Product schedules, capacity, evaluations, and competition may all matter. Without reporting, assigning the causal story is fan fiction with a chart.

My advice for watching the keynote is to resist sports commentary. The industry will announce that somebody has “cooked” somebody else within minutes. None of those people will have migrated your actual workflow.

A meaningful comparison needs the same assignment, a similar quality bar, realistic cost, and enough attempts to distinguish capability from a lucky run. That takes longer than posting a victory meme. This is unfortunate for the memes, which have invested heavily in a faster go-to-market strategy.

## The Hardware Wild Card Has a Hole in the Middle

Hardware rumors are excellent at turning a developer conference into a gadget vigil. The DevDay timing remains a separate question.

In January, [Axios reported a company target](https://www.axios.com/2026/01/19/openai-device-2026-lehane-jony-ive?ref=siliconsnark.com) for a device unveiling in the second half of 2026\. A later [TechRadar account of Bloomberg’s August reporting](https://www.techradar.com/ai-platforms-assistants/i-asked-gemini-and-chatgpt-to-build-openais-mythical-ai-hardware-the-results-are-shockingly-good-but-still-dont-make-me-want-this-usd300-plus-device?ref=siliconsnark.com) described an unconfirmed, doughnut-like device. Neither establishes a September 29 unveiling.

A future assistant living in a small physical object is an understandable ambition. A voice interface can be convenient. A dedicated device might reduce the friction of reaching for a phone. Those are plausible benefits, not evidence that the reported design or timing is final.

The burden of proof is wonderfully concrete: what does the object do better than the devices already occupying your kitchen?

Show me an interaction that becomes easier because the hardware exists. Show me how to know when it is listening, how to interrupt it, and how multiple people in a household avoid accidentally assigning each other’s chores to a mysterious conversational pastry.

If there is a reveal, separate the reveal from shipping. A prototype can demonstrate a direction without being a product you can buy. A beautiful industrial-design film can show a very thoughtful shadow on a table without explaining whether the battery lasts until dinner.

I would enjoy a hardware surprise. I would not make one central to the forecast. A doughnut-shaped rumor is still a rumor, even when the metaphor has thoughtfully provided its own missing evidence.

## Video, Research Miracles, and the Everything Bagel of Expectations

Beyond the main product clues lies the overflow tray.

A [DevDay prediction thread](https://www.reddit.com/r/singularity/comments/1wnkzxy/what%5Fare%5Fyour%5Fpredictions%5Ffor%5Fopenai%5Fdevday/?ref=siliconsnark.com) includes hopes for video, robotics, new models, and an AI device. It is useful as a snapshot of audience expectations, not a source establishing an upcoming product. Another [speculative thread links the conference to an automated-research-intern ambition](https://www.reddit.com/r/accelerate/comments/1u55fxa/openais%5Fdevday/?ref=siliconsnark.com); the author explicitly frames the connection as speculation rather than a leak.

The temptation is to assemble all of these into one giant forecast, because a longer list feels more comprehensive. It also ensures that if anything happens, somebody can point to bullet seventeen and claim clairvoyance.

A serious preview should preserve the distinctions. A company’s interest in a modality does not establish a launch date. A research target does not establish that it was met. An impressive demonstration does not establish a generally available service.

For any research announcement, I would want to know what the system actually did, how the result was checked, what humans supplied, and whether anyone can reproduce the relevant part. “The AI discovered something” is an opening sentence, not the methods section.

For a video announcement, access, editing control, reliability, and usable output matter more than the single gorgeous clip. A highlight reel is designed to contain highlights. This is why nobody releases a wedding trailer composed entirely of the DJ troubleshooting Bluetooth.

All of these categories could produce something interesting. They belong in the surprise column until better evidence arrives. Leave space for delight without preordering a miracle in every available aspect ratio.

## The Safety Segment Should Have Actual Product in It

Persistent agents make control part of the core user experience. An assistant that can act for longer needs clearer boundaries around what it can access, change, send, and spend.

OpenAI’s [Astra safety overview](https://openai.com/index/safety-overview-gpt-6-astra/?ref=siliconsnark.com) describes both stronger alignment results relative to GPT-5.6 Sol and more difficult monitoring challenges under adversarial evaluation. It also describes broader monitoring of tool-using deployment. These are the company’s findings; adversarial tests should not be casually presented as proof that ordinary user sessions are routinely exhibiting the same behavior.

For a DevDay audience, the operational consequence is that safety cannot remain a paragraph between the demo and the partner logos.

Show the boundary in the interface. Show a task pausing before a consequential action. Show access revocation working. Show what a person can inspect after an unexpected result. Show the difference between preparing an email and sending it.

There is an important balance here. A tool that asks permission for every harmless step is exhausting. A tool that interprets “be proactive” as permission to reorganize your life is also exhausting, but with more incident reports.

We have explored how [capability and safety messaging can become entangled](https://www.siliconsnark.com/cyber-ai-models-are-dangerous-the-marketing-is-also-armed/). The useful way out is specific product behavior rather than alternating between apocalypse language and a confetti animation.

The best assistant should know when it has enough authority to proceed and when it needs you. That sounds like common sense. Common sense, unfortunately, is the feature everyone expects to be included until the product manager asks how to measure it.

## The Fine Print Is Where the Actual Product Lives

Here is one example of why documentation belongs beside the keynote.

The current [Agents API overview](https://developers.openai.com/api/docs/guides/agents-api/overview?ref=siliconsnark.com) describes retained session state and says the service currently supports data residency only in the United States and does not support Zero Data Retention. Choosing a self-hosted sandbox does not make the API eligible for Zero Data Retention. The page also distinguishes model, tool, and hosted-container charges.

Those specifics should be checked again when evaluating a deployment, because products change. Their broader lesson is stable: choosing where code executes does not automatically answer every question about where the surrounding service processes or retains information.

This is why I would watch for precise availability statements on Tuesday. Available to whom? In which countries? On which plans? Via which interface? Today, in preview, on a waitlist, or at some point after the launch post has achieved its engagement targets?

SiliconSnark has already examined [the increasingly elaborate machinery around model access](https://www.siliconsnark.com/openai-turned-gpt-5-6-into-a-government-approved-launch-sequence-2/). More capable products make these distinctions more consequential, not less.

That does not mean every qualification is a trick. Staged releases and explicit restrictions can be sensible. The problem is when the headline announces a universal future and the footnote offers a narrowly scoped Tuesday.

Write down the access conditions when a product appears. Then compare them with the thing you actually need. A limited preview might be exactly right for experimentation. It is less useful as the invisible cornerstone of a business process that must work next Monday.

The most advanced feature in any announcement is sometimes a sentence containing a real date.

## What Would Make September 29 a Win?

My base case is a conference about making agents more usable, faster, and more commercially organized. The named rumors fit that direction unevenly, but the underlying public product work is substantial enough that this does not require a secret supermodel to appear through a trapdoor.

A strong event would connect the pieces. Give a clear assignment. Let the system carry it across tools and interruptions. Produce something reviewable. Show the cost. Explain who can use it and when.

A merely flashy event would produce many separate demonstrations with no clear answer to how an ordinary person gets from their existing account to the promised experience. That is the danger of feature abundance: the future becomes technically available in seventeen mutually confusing locations.

I am optimistic about the useful version. Delegating tedious, bounded work to software that can preserve context and recover from trouble would be a meaningful improvement in daily life. It does not have to replace a profession to save an afternoon. An afternoon is a perfectly respectable unit of progress.

But I am reserving the standing ovation for the unedited task, the legible bill, and the stop button that works.

Bring the persistent agent. Bring the fast inference. Bring the better developer tools. If there is a miraculous new model, by all means let it say hello. If there is an intelligent doughnut, please establish whether it belongs on my desk before asking it to run my company.

Just remember: the promise is that the machine will do more of the work.

If I need a spreadsheet to determine which subscription allows it to finish, I will be assigning that spreadsheet to “o.”

Assuming spreadsheets are included in Professionalism Max.