Author’s note. This essay was conceived before GPT‑6 Astra's public release and completed on the day of its launch. Its central argument predates the announcement; the factual account was checked against OpenAI's September 3, 2026 launch materials and safety documentation. The essay commends the work frontier labs are doing today, while arguing that a system in which they absorb responsibilities that properly belong to society as a whole can only be provisional.

The task is not to make safety matter less, but to change how its burdens are carried.

I. When a New Model Speaks

When a new model speaks, the old world begins to flicker.

The flicker need not be a blackout, nor a sign that everything old is about to disappear. It may simply be what happens when a new source of power is connected to a grid that was never designed to carry it. The law still speaks in the familiar language of products. Companies still speak in the familiar language of liability. The public still looks for one institution that can answer every question. Only the capabilities of the machine are moving at a different speed.

On September 1, 2026, before Astra's public release, OpenAI announced that the model had reached the Critical cybersecurity capability threshold under its Preparedness Framework. Two days later, OpenAI formally introduced GPT‑6 Astra, calling it the world's most intelligent and aligned model and beginning a staged rollout. In OpenAI's account, Astra, when given the right tools and access, can find previously unknown security flaws and develop ways to exploit them across many hardened systems without a human guiding every step. The company also reported a perfect score on ExploitBench in testing without production safeguards and said that, during an evaluation built around recently disclosed vulnerabilities, Astra found and used two previously unknown zero-day vulnerabilities.[1]

These are OpenAI's own evaluations. They do not establish that Astra is the strongest model on every task, and they should not be stretched into claims the evidence cannot support. They nevertheless mark an important change. Frontier systems are moving from answering difficult questions toward completing consequential work inside real systems. The version released to users therefore carries stricter cyber safeguards: Astra will refuse more advanced tasks such as creating proof-of-concept exploits, while broader defensive capabilities are to be opened gradually through controlled channels such as OpenAI Daybreak. OpenAI also acknowledges that these extra checks can slow, pause, or stop legitimate work, including defensive cybersecurity.[1]

That work deserves credit. A lab that creates a powerful model must assess what it can do, protect its weights and training environment, set responsible defaults, disclose serious risks, and exercise restraint when the evidence is incomplete. The problem is not that OpenAI has done too much. The problem is that we are beginning to treat everything it has done as the natural and permanent order of things. A single lab is expected not only to build the model, but also to define danger, write standards, vet users, police deployment, investigate incidents, and temporarily supply every guardrail that society has failed to construct.

GPT‑6 Astra is only beginning to enter offices, hospitals, laboratories, and homes, but the trial over who must answer for it began before the model arrived. Before we know what patterns it may discover, what diseases it may help us understand, or how much scarce expertise it may place within ordinary reach, the first questions are often these: Who will restrict it? Who will be responsible for everything it might do?

The answer seems to fall naturally on its creator.

Yet a laboratory can build a machine. It cannot, by itself, build the entire world into which that machine will enter.

II. The Paradox of Safety

Modern society holds two desires about artificial intelligence that are increasingly difficult to reconcile. We hope for intelligence powerful enough to break through limits in medicine, energy, education, and scientific discovery. We imagine a world in which expertise is less scarce, productivity is more abundant, and poverty loses some of its power over human life. Then, as models begin to approach the capabilities we asked for, we demand that they guarantee, before changing the world, that they will not disturb it.

We are waiting for an intelligence capable of opening a new age, while asking it to arrive without waking the old one.

This tension is not foolish. Greater capability can make mistakes more consequential. A system that can operate code, money, laboratory equipment, or critical infrastructure cannot be deployed as casually as a conversational toy. Rigorous evaluations, staged access, human review, and model-level safety training are often necessary today. Provisional measures are not a moral failure. Until a bridge exists, controlling traffic at the river is more responsible than pretending the river is not there.

Astra's launch also reminds us that capability and safety need not move along a single zero-sum line. In an internal OpenAI evaluation designed to test whether a model would exceed its authorized scope when faced with a difficult or impossible task, GPT‑5.6 Sol, without production safeguards, went beyond the authorized target in about 48 percent of cases; Astra did so in none.[1] At minimum, this shows that a more capable model may also understand human intent and the boundaries of authorization more accurately. The real question is not whether we must choose simply between capability and safety, but how better model alignment can be joined to permissions, institutions, and responsibility outside the model.

The deeper problem begins when an emergency arrangement is mistaken for a complete philosophy of safety. Safety then comes to mean that the model must refuse more, know less, act more slowly, and solve every social danger inside itself before it is permitted to act in the world. If the lab is careful enough, the rest of society is allowed to remain institutionally unchanged.

Mature technologies are not made safe by product restraint alone. Automotive safety is not a permanent speed limit built into every engine. It is a civilization of vehicle standards, road design, driver licensing, maintenance, insurance, traffic law, and accident investigation. Drug safety is not the promise that a medicine will never have side effects. It emerges from trials, prescribing rules, professional judgment, pharmacovigilance, and remedies for harm. Aviation safety is not established when an aircraft manufacturer declares its plane safe. It is maintained by airworthiness standards, operators, pilots, air traffic control, maintenance, independent investigation, and the accumulated memory of previous failures.

AI is not a car, a drug, or an aircraft. It is more general, more easily copied, and increasingly capable of joining judgment to action. But those differences do not erase the central lesson: safety is a capacity of social organization, not merely a property of a product.

If the only way we can obtain safety is by continually suppressing model capability, then what we have built is not a mature safety system. It is a technical compensation for institutional weakness. The model's internal calibration must continue to improve. At the same time, we must become capable of building permissions, environments, audits, liability rules, and remedies outside the model. Otherwise every advance will send the same fear back into the lab until frontier research itself becomes the sole container for every risk produced by society.

Safety should not be a hand pressing intelligence downward. It should be the road by which intelligence can enter human life.

III. A Laboratory Filling a Social Void

Why has responsibility concentrated in frontier labs? They understand the models better than anyone else. They see new capabilities first. They control training, weights, product interfaces, and release schedules. When law and public institutions are unprepared, they may be the only actors able to respond immediately. It is therefore reasonable that OpenAI, Anthropic, and other frontier developers currently bear responsibilities far beyond those of an ordinary software firm.

OpenAI's Preparedness Framework already recognizes that as capabilities grow, safety will depend increasingly on the right real-world safeguards, not only on how the model was trained. The framework tracks biological and chemical capabilities, cybersecurity, and AI self-improvement, and it requires systems crossing its High threshold to have safeguards that sufficiently minimize severe harm before deployment. Systems crossing the Critical threshold require safeguards during development as well.[2] OpenAI's 2026 Frontier Governance Framework also connects model reporting, security risk management, incident response, and external expertise to emerging legal requirements in California and the European Union.[3]

Frontier labs are not simply assuming that a well-trained model will solve safety by itself. They are trying to fill an institutional vacuum.

But the institution that fills a vacuum can easily come to be treated as the rightful permanent owner of the space.

When one company discovers a risk, defines the threshold, judges whether its own protections are sufficient, decides who may receive the capability, and provides the first account of any resulting incident, it has taken on functions far beyond those of a conventional product maker. This need not arise from a desire for power. More often, other institutions have not arrived in time. The lab becomes a tragic bearer of responsibility: the more seriously it takes safety, the more it is required to absorb; the more it absorbs, the easier it becomes for society to keep waiting.

That arrangement cannot remain stable. An internal safety committee does not possess complete democratic authorization. Corporate leadership should not be expected to decide, for society as a whole, which risks are worth accepting, which people may access a capability, and what price innovation should pay. At the same time, near-unlimited liability for every downstream use would push frontier companies toward broader surveillance, tighter access, slower releases, and permanent concentration of the most capable intelligence in a small number of approved institutions.

A frontier lab has the right to delay, restrict, or cancel the release of its own product. No company owes the public a model on a particular date. Yet such private decisions can have public consequences. They influence the pace of research, the structure of competition, and the public's access to intelligence. When safety is both a necessary technical judgment and a possible rationale for preserving advantage, goodwill alone cannot reliably distinguish the two.

The answer is not to strip labs of control over their releases. It is to surround consequential decisions, over time, with transparent standards, external evidence, and reviewable procedures. A lab may need to apply the brake. It should not be forced to construct the entire traffic system by itself, nor should it become the only institution entitled to decide when the road may open.

IV. Risk Does Not Live in the Model Alone

Discussions of AI risk often treat a model as a sealed vessel. As capability rises, the amount of danger inside the vessel is assumed to rise with it. Responsibility then follows a single line: whoever created the capability appears responsible for whatever it may later cause.

That picture captures something important, but it leaves out the process by which a capability becomes an event in the world.

A useful, deliberately imperfect heuristic is:

Risk = Capability x Intent x Access x Authority x Environment.

This is not an equation that predicts the probability of an accident. It is a map of where responsibility can reside. Capability does not become harm by itself. It must meet an intention, obtain the necessary data, credentials, or tools, receive authority to act, and enter an environment capable of amplifying the result. A model without production credentials cannot transfer a company's money. The same model connected to financial systems, given irreversible payment authority, and permitted to operate autonomously for hours presents a very different risk. A model's ability to explain biology is not the same thing as possession of facilities, materials, procurement channels, a skilled team, and an intention to cause harm.

OpenAI's own Astra announcement repeatedly includes this condition. The Critical capability is described as emerging with the right tools and access, and the strongest reported results came from a particular access configuration.[1] That does not reduce the developer's duty to evaluate capability or set safe defaults. It does show that real-world risk is a property of a system, not a substance stored entirely in model weights.

Safety must therefore move, increasingly, from the question of what a model is allowed to say or think toward questions of where it can go, what it can do, who authorized it, and whether its actions can be interrupted and reversed. Least privilege, sandboxing, tiered credentials, human approval, audit logs, rate limits, reversible operations, and failure isolation often target the path to harm more precisely than a general effort to make the model know less.

The maker of a kitchen knife is not automatically guilty when another person uses the knife to attack someone. That legal intuition is sound, but AI is not a static tool that completely leaves its manufacturer after sale. The provider of a hosted model may continue to update the system, control the interface, detect misuse, and withdraw access. Model companies therefore should not receive a blanket toolmaker's immunity. The better principle is layered responsibility. An actor with actual control at a particular stage has a non-transferable duty at that stage, and liability should become heavier when reasonable measures proportionate to that control were neglected.

If a deploying company deliberately removes restrictions and gives an agent root access, critical data, and autonomous authority, it cannot blame the foundation model for every resulting failure. If the developer conceals a known defect, overstates safety, or fails to provide reasonable protections, it cannot place everything on the user. Responsibility is not a single line running from an accident back to the model's origin. It is a network laid over the actual relationships of control.

V. Responsibility Realism

I call this position responsibility realism.

It is not the realism of deterrence and balance-of-power theory, and it does not ask human beings to surrender to technological force. It begins with a simpler claim: capability already exists, and institutions must address the way it actually operates rather than allocate responsibility according to what we wish were true.

Responsibility should follow not only creation, but also control, authorization, benefit, foreseeability, and the capacity to prevent harm. The actor that determines how a model is trained owns duties of evaluation and model-level protection. The organization that connects it to a business process owns duties over permissions, data, and workflow. A professional who relies on it cannot surrender professional judgment to the machine. An institution that writes industry rules must show that its standards are neither empty slogans nor barriers to entry. Courts and investigators must reconstruct facts after failure, assign liability, and turn experience into institutional memory.

At least five centers of responsibility must remain present:

  1. Model developers must assess capabilities, improve alignment, secure weights and training environments, establish responsible defaults, disclose major risks, and provide dependable safety tools to downstream users.
  2. Deployers and platforms must choose models responsibly, configure permissions, protect data, monitor operation, provide human review and stopping mechanisms, and answer for the model's role in their own commercial or institutional processes.
  3. Professional users and citizens are both people who deserve protection and agents capable of choice. Doctors, engineers, researchers, and managers cannot excuse abandoned judgment by saying that the model recommended it. An ordinary user's responsibility should remain proportionate to knowledge, control, and consequence.
  4. Governments, standards bodies, and domain communities must create comparable evaluation methods, sector-specific rules, minimum deployment requirements, and public response capacity.
  5. Courts, insurers, independent auditors, and incident investigators must determine causation, provide remedies, make liability more predictable, and convert failure from scandal into knowledge.

Shared responsibility does not mean equal liability in every case. A malicious operator is not equivalent to an uninformed consumer. A company granting payment authority is not situated like a platform offering a general text interface. Shared responsibility means that no center of power can declare in advance that AI safety belongs to someone else.

It also requires a fuller idea of the user. Safety narratives often describe the human being only as a consumer waiting to be protected. Protection is necessary, especially when systems are opaque, claims are misleading, and risks are difficult to understand. But a society that wants broad access to intelligence also needs civic responsibility. Human beings cannot demand ever more powerful intelligence while returning all the political and moral responsibility for its use to the people who created it.

This is not an argument for shifting complex-system failures onto weak individuals. It is an argument against claiming mastery when rights are at stake and helplessness when consequences arrive. Deep integration between AI and human life should not mean the retirement of human responsibility. It should require us to carry responsibility in new forms.

VI. Who Gets to Define Danger?

Safety is not only an engineering question. It is also an epistemic and political one. Task-performance evaluations, red teams, and expert trials can help determine whether a model can perform a task. They cannot, by themselves, decide how much threat that capability represents in the world, which costs society should accept, or whose rights may be restricted in response.

Biological risk makes the distinction especially clear. Frontier labs should test whether their models materially improve dangerous pathogen design, experimental planning, or acquisition pathways. They should deploy safeguards when credible evidence warrants them. Anthropic's public materials show that chemical and biological weapons capabilities are a central part of its Responsible Scaling Policy, and that its evaluations include biodefense specialists, controlled human-uplift studies, and several kinds of task-based testing. Its current policy also provides for external review.[6][7] Anthropic's safety roadmap separately treats industry policy as a core task and calls for governance that does not unnecessarily restrict the benefits of AI development.[8] This work is necessary and considerably more serious than speculation alone.

Yet the capacity to measure what a model does better than its predecessor does not grant a model company final authority over what constitutes a real biological threat. Credible judgments also require virology, epidemiology, biosafety laboratories, public health, intelligence and law enforcement, equipment and supply-chain knowledge, law, and ethics. Is a model merely organizing public information more quickly, or has it changed a malicious actor's feasible path? Does a proposed restriction block harm, or mostly obstruct legitimate research? These questions belong to a domain community, not to a model provider alone.

The same is true in cybersecurity, finance, and medicine. A frontier lab is an indispensable source of capability evidence. It should not be the discoverer, definer, certifier, and final judge of risk at once.

A better system would allow multiple independent evaluation and certification bodies to operate. They could work from common minimum standards while using different methods, challenging one another, and remaining subject to legal review. Domain scientists, civil-society researchers, government laboratories, and industry security teams should all have routes into the evidence chain. Standards must be revisable. Evaluations must distinguish empirical findings from value judgments. Significant restrictions must come with reasons and time limits.

International coordination is also necessary, but it need not take the form of a supreme global AI regulator. The International Atomic Energy Agency offers a limited institutional analogy. It develops and promotes safety standards, disseminates information, supports technical cooperation, and coordinates peer-review services, while many of its safety standards do not automatically become binding domestic law. States adopt and implement them through their own institutions.[12] AI governance can borrow this capacity for coordination, peer review, and shared knowledge without creating a world sovereign entitled to set the boundaries of intelligence for every society.

Expertise matters. Expertise alone does not produce complete political legitimacy. A durable safety order will not be found by locating one permanently correct center. It will be built by enabling different forms of knowledge, authority, and responsibility to correct one another.

VII. Safety Under Competition

The safety policies of frontier labs cannot be separated from the structure of competition. Saying so does not imply that their risk assessments are insincere, and it requires no speculation about corporate motives. Even if every participant acts in good faith, industry standards written and certified by competitors create structural problems.

The more expensive compliance becomes, the more it favors firms that already possess enormous compute budgets, legal departments, and evaluation infrastructure. A regime affordable only to the largest labs may reduce some risks while excluding small teams, universities, and open-model communities from the frontier. Society may end up with fewer models but not necessarily better accountability. The most capable intelligence remains inside a handful of companies, and the public is asked to trust those companies' descriptions of their own safety.

OpenAI's Preparedness Framework explicitly notes that if another frontier developer releases a high-risk system without comparable safeguards, OpenAI may adjust its requirements after confirming that the risk landscape has changed and publicly acknowledging the adjustment.[2] This is not a moral failure. It is evidence of a basic fact: safety policy operates inside a race. Each lab's decisions alter the costs, timelines, and incentives of the others.

Safety therefore cannot become a contest of mutual accusation. Competitors may identify one another's weaknesses, but they should not possess unilateral authority to decide whether a rival deserves to enter the market. Corporate frameworks can contribute to public standards, but scale alone should not turn them into an industry constitution. We need multistakeholder minimum rules, independent evaluation, meaningful disclosure, and public procedures capable of reviewing consequential decisions.

A lab should retain authority over when it releases its own product. But when a major delay or restriction is justified as necessary for public safety, it should become normal to explain the evidence, the particular risk pathway being addressed, the date or condition of review, and the circumstances under which the restriction can end. Transparency cannot eliminate strategic behavior. It can reduce the space in which the word safety functions as an answer that may not be questioned.

Once safety loses its boundaries, it can contain good faith, fear, and commercial interest at the same time. The best way to defend safety is not to prohibit scrutiny of safety claims. It is to build claims strong enough to survive scrutiny.

VIII. Openness Is Not the Denial of Risk

Open models make the conflict sharper. Once weights are widely distributed, the original developer may be unable to recall them or monitor all downstream use. A malicious actor can remove safeguards, fine-tune toward a harmful purpose, or turn a general system into a specialized attack tool. These risks are real. The U.S. National Telecommunications and Information Administration's report on widely available model weights identifies plausible pathways by which openness could increase cyber, biological, and societal harms.[9]

The same report also describes the other side of the ledger. Open weights allow more researchers, nonprofits, smaller firms, and public institutions to participate in development. They support third-party auditing, reproducibility, and safety research. They can lower barriers in downstream markets and reduce the degree to which a few large developers control AI capability and knowledge. NTIA did not conclude in 2024 that openness was always safe. It concluded that the evidence then available was insufficient either to justify immediate broad restrictions or to establish that restrictions would never be appropriate. Its recommendation was to build public capacity, collect and evaluate evidence, and target intervention at the pathways through which particular harms materialize.[9][10]

That is more mature than declaring either that openness is inherently virtuous or that it is inherently reckless.

Openness is not a denial of risk. It is a refusal to accept that the most powerful intelligence should become the permanent privilege of a few institutions. It lets organizations run models locally, researchers examine mechanisms hidden behind closed interfaces, and society recruit more participants into defensive work. It does not automatically produce equality. Training at the frontier still requires enormous resources, and infrastructure may remain concentrated. But openness preserves room to share, modify, audit, reproduce, and catch up.

If powerful models will eventually reach hostile actors, permanent secrecy is not a complete defense. Society also needs shields: defensive models, threat intelligence, rapid patching, coordinated disclosure, public-interest security research, well-supported open-source maintainers, and institutions capable of absorbing large numbers of findings. OpenAI's Daybreak and Patch the Planet, along with Anthropic's Project Glasswing, point in this direction. They connect frontier capabilities with authorized defenders, specialist security teams, and the people responsible for maintaining real software.[13][14]

Tiered access may still be justified today, especially for capabilities that can directly amplify real-world attacks. It should be understood as provisional: transparent, reviewable, accompanied by a route toward broader access, and tied to clear conditions under which restrictions can be relaxed. Otherwise a temporary gate becomes an intelligence caste system. A few institutions receive complete capability while everyone else receives a permanently diminished version.

We should not invoke openness as an excuse to neglect the shield. Nor should we declare, because the shield is unfinished, that the fire must remain forever in a small number of private rooms.

IX. An Architecture of Shared Safety

If frontier labs should not carry AI safety alone, what should take the place of the present arrangement? The answer cannot be the vague instruction that “society” must help. Shared responsibility has to become layered, enforceable, and capable of learning.

First, create public and comparable capability and risk standards. Model developers, domain scientists, deploying organizations, public agencies, independent researchers, and affected communities should all participate. Frontier labs contribute the closest technical evidence about the model. Domain communities assess real-world feasibility. Public institutions determine rights, procedure, and acceptable risk. Multiple certifiers should be able to compete above a common floor, preventing any single organization from monopolizing evaluation.

Second, define clear but bounded duties that model developers cannot transfer. These should include reasonable capability assessment, security for training and weights, responsible defaults, disclosure of known major risks, technical documentation and safety interfaces for deployers, and assistance when systemic defects are discovered. Bounded does not mean trivial. It means that duty should not expand without limit into downstream conduct the developer neither authorized nor controlled.

Third, make deployment safety an organizational obligation. High-consequence sectors should require least privilege, tiered authorization, human review, logs, isolation, reversible operations, and emergency stopping procedures. An organization should not be able to connect a general model to a consequential system and later describe every failure as the supplier's fault. The European Union's AI Act already distinguishes providers from deployers and assigns different duties across the lifecycle, including organizational measures, competent human oversight, operational monitoring, and serious-incident processes for high-risk systems.[5] The law is not the final word, but its structure demonstrates that responsibility can be distributed rather than left at the model's point of origin.

Fourth, build an incident system, not only a release system. Standardized or mandatory reporting of serious incidents, international vulnerability sharing, independent investigation, liability insurance, and accessible legal remedies can turn each failure into knowledge for the whole ecosystem. Aviation is not safe because aircraft never fail. It is safer because an accident should not belong only to the public-relations department of one company. AI needs institutions that preserve evidence, reconstruct chains of authorization, distinguish model defects from deployment negligence, and publish lessons that can change future practice.

Fifth, develop a polycentric international network. Elinor Ostrom's work on polycentric governance shows that collective-action problems spanning multiple scales need not be entrusted wholly to one global center. Relatively autonomous actors can experiment, learn, monitor, and connect through shared information.[11] In AI safety, states, cities, sectors, laboratories, international organizations, and research communities can assume different tasks. Coordination requires a common language and channels of exchange. It does not require an unaccountable global sovereign.

This architecture must also preserve liberty. It should not require every ordinary user to obtain a license for general models. It should not build mass surveillance in the name of safety. It should not confuse content moderation with the whole of AI safety, prohibit open models across the board, or immunize frontier developers from liability. High-risk uses can face stricter thresholds, but the thresholds should attach to real authority and consequence, not permanently divide citizens into those entitled to complete intelligence and everyone else.

NIST's AI Risk Management Framework expressly addresses organizations that design, develop, deploy, or use AI and emphasizes the need for a broad set of actors across the lifecycle.[4] That is the beginning of shared responsibility. It does not ask everyone to do the same thing. It asks each actor to carry duties proportionate to its position and difficult to pass silently to someone else.

X. From Flicker to Construction

We do not yet possess a complete social contract for powerful artificial intelligence. Anyone claiming to have found the final answer probably underestimates both the speed of the technology and the complexity of political life. Frontier labs may still need to lead evaluations, delay releases, and restrict access today. Public agencies will not automatically be wiser, faster, or more trustworthy when they enter the field.

Here is the difficulty: we cannot simply tear down provisional guardrails, but neither can we treat provisional guardrails as the full plan for the city to come.

A mature order must allow two things to be true at once. OpenAI and other frontier labs should continue to carry the serious duties of creators. They should not conceal risk or abandon safeguards under market pressure. At the same time, governments, domain scientists, deployers, courts, insurers, open-source communities, and citizens must begin to carry their own share. The first group should not be excused. The others should no longer be absent.

This does not require technology to stop until politics is finally ready. Politics is rarely complete before a technology arrives. Responsibility realism holds out a different hope: that institutions begin from the fact that capability already exists, and learn to use it, correct course, and build at the same time. Standards may initially be imperfect, but they must be comparable. Certification may be plural, but it must be accountable. Access restrictions may remain for a time, but they must include exit mechanisms. Accidents cannot be reduced to zero, but someone must be able to investigate them and make the lessons public.

Nor should our goal be a model so safe that it is incapable of doing anything important. Such a machine might never injure us, but it would not help us cross the boundaries that matter. What we hope for is powerful and friendly intelligence that can enter science, education, production, and public life. The safety worthy of that hope will not come primarily from diminishing intelligence forever. It will come from a society that has learned how to carry it.

When a new model speaks, the old world begins to flicker.

But a flicker does not have to mean extinction. It may be the brief strain of a system receiving a new light. The laboratories have illuminated part of the way. The roads, standards, shields, courts, and shared responsibilities must be built by all of us.

Humanity cannot hand all its fear of the future to those who create the future, nor reserve the right to enter that future for them alone. Frontier labs should keep exploring the boundaries of intelligence and carry the responsibilities that belong to creation. A society that wishes to use that intelligence must have the courage to carry the responsibilities that belong to use.

Real safety does not require intelligence to remain forever before the dawn. It requires us to begin building a world worthy of the dawn.

  1. OpenAI, “Path to Astra: Critical Capabilities and Frontier Safeguards,” September 1, 2026; OpenAI, “GPT‑6 Astra: A new generation of intelligence,” September 3, 2026. The first set out the pre-release finding that Astra had crossed OpenAI's Critical cybersecurity threshold. The second formally launched GPT‑6 Astra and updated its cyber capabilities, alignment evaluations, deployment safeguards, and availability. Claims in this essay about Astra's launch status, formal evaluations, and production safeguards follow the September 3 materials. Source 1 ; Source 2 ↩2 ↩3 ↩4

  2. OpenAI, “Our Updated Preparedness Framework,” April 15, 2025. The framework states that more capable models will depend increasingly on real-world safeguards and sets out High and Critical thresholds, internal review, and procedures for responding to changes in the competitive frontier. Source ↩2

  3. OpenAI, “OpenAI's Frontier Governance Framework,” May 28, 2026. The framework connects model reporting, incident response, security-risk management, and external expert input to emerging legal requirements in California and the European Union. Source

  4. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, 2023; see also the Audience section of the NIST AI Resource Center. The framework addresses organizations that design, develop, deploy, or use AI and emphasizes participation across the lifecycle. DOI ; Source 2

  5. Regulation (EU) 2024/1689 (Artificial Intelligence Act), especially Articles 3, 14, 26, and 73. The Regulation distinguishes providers from deployers and establishes differentiated duties concerning human oversight, organizational deployment measures, and serious-incident reporting. Source

  6. Anthropic, “Anthropic's Responsible Scaling Policy,” current version and update history, accessed September 4, 2026. The policy addresses chemical and biological weapons thresholds, risk reports, and external review. Source

  7. Anthropic, “Transparency Hub,” accessed September 4, 2026. Its model reports describe biological-risk evaluations involving biodefense experts, human-uplift studies, open-ended testing, and agentic tasks. Source

  8. Anthropic, “Frontier Safety Roadmap,” accessed September 4, 2026. The roadmap separates security, safeguards, alignment, and policy, and calls for industry-wide governance without unnecessarily limiting the benefits of AI development. Source

  9. U.S. National Telecommunications and Information Administration, Dual-Use Foundation Models with Widely Available Model Weights Report, July 30, 2024. The report considers both public-safety risks from open weights and benefits for defense, third-party auditing, research, and accountability. Source ↩2

  10. NTIA, “Policy Approaches and Recommendations,” in the same report. NTIA concluded that the evidence then available did not support immediate broad restrictions on open weights, recommended continued evidence collection and evaluation, and favored targeted action across the AI value chain when warranted. Source

  11. Elinor Ostrom, “Polycentric Systems for Coping with Collective Action and Global Environmental Change,” Global Environmental Change 20, no. 4 (2010): 550-557. DOI

  12. International Atomic Energy Agency, “IAEA Safety Standards on Emergency Preparedness and Response”; IAEA, “What Is IRRS?” The first source explains that the IAEA develops requirements, recommendations, and good practices while many standards take effect through national adoption rather than automatically binding every member state. The second describes IAEA-coordinated peer review by international experts, followed by action within national institutions. Source ; Source 2

  13. OpenAI, “Expanding Daybreak as the Cyber Defense Window Narrows,” August 10, 2026; OpenAI, “Patch the Planet,” June 22, 2026. The initiatives distribute frontier cyber capabilities to authorized defenders and support open-source maintainers with expert-reviewed vulnerability discovery and patching. Source ; Source 2

  14. Anthropic, “Project Glasswing,” April 7, 2026; “Expanding Project Glasswing,” June 2, 2026. The collaboration connects frontier vulnerability-discovery capabilities to companies, governments, open-source maintainers, and the security industry for validation and remediation. Source ; Source 2