Governance

Rules, incentives, and institutions for AI that stays under human control and serves human interests.

The Pro-Human Tool Framework describes what we should build. Governance is how we make it happen: creating conditions where controllable, trustworthy, pro-human AI is the path of least resistance, and where dangerous paths are foreclosed.

Governance

What This Section Is

This is a proposal for the governance we need, not a description of what is politically feasible at this moment. We expect the importance of AI governance to increase dramatically in the coming months as AI systems become more capable and their effects become harder to ignore. When that window opens, the response should be guided by a coherent framework prepared in advance, not assembled from piecemeal half-steps.

What follows is such a framework. It describes six areas of governance, each addressing a distinct part of the problem, designed to work as an interlocking system. We believe that even if imperfect, if enacted, this framework would largely transition us to the Pro-Human Path.

These are not draft laws. They describe what governance should accomplish, the mechanisms and incentive structures that can accomplish it, and how the pieces fit together. Any given component could be implemented through different legislative vehicles, regulatory structures, or international instruments depending on jurisdiction and political context. Given the right substance, the specific legal architecture follows. That said, FLI does have sample legal text implementing many parts of this framework; please get in touch if you're interested.

The Governance Challenge

AI governance must address the following problems simultaneously, because they are largely interdependent. Solving any one without the others leaves critical gaps.

Runaway uncontrollable systems. The corporate race toward AGI and superintelligence, amplified by geopolitical rivalry, pushes development toward systems too powerful and autonomous to meaningfully control. Without hard limits on capability scaling and coordination to prevent race dynamics, these systems are built in a way virtually guaranteeing that humans won't be in charge afterwards.

Large-scale and catastrophic risk. Even short of superintelligence, highly capable systems – whether through misuse, accident, or for their own purposes – can cause harm ranging all the way up to societal or even civilization-level if developed unsafely and managed poorly. And on an everyday level, when AI systems make consequential decisions without meaningful human oversight, the question of who bears responsibility becomes urgent. Without clear answers, the default is that no one does.

Pervasive everyday harms. AI systems are already manipulating attention, distorting information, forming engineered relationships, and eroding privacy. These harms are diffuse and cumulative. They don't require dramatic failures; they emerge from systems working exactly as designed, optimizing for engagement and profiting at the expense of humans and society.

Concentration of power. AI development is extraordinarily concentrated. A handful of companies control the most capable systems, the compute infrastructure, the data pipelines, and increasingly the platforms through which AI reaches users. Without countermeasures, this concentration deepens, creating private power that is difficult for democratic institutions to check.

Capture benefits, as well as manage the risks. AI, as a powerful tool, can be used for enormous good and to address many of the world's problems. Governing it well means not just avoiding the risk, but putting in place the right incentives for it to quickly and broadly benefit humanity.

Placeholder

Argonne National Laboratory staff members are demonstrating the type of purpose-driven, bounded AI that complies with a robust governance framework. This experiment connects a beamline with a supercomuter for rapid data analysis with real-time feedback for researchers to execute quicker and more refined experiments. This approach leverages transparent, human-led verification and rigorous assurances for AI without sacrificing accelerated scientific discovery. (Image by Argonne National Laboratory.)

The Approach

This goal of this governance framework is to redirect AI development from the Race to Replace to the Pro-Human Path. We want conditions where beneficial AI flourishes, while dangerous paths are closed off.

Clear liability rules would reduce the legal uncertainty that currently discourages responsible deployment, and safe harbors give developers who build controllable systems a concrete compliance advantage. Certification can create warranted trust that unlocks adoption in high-stakes settings where unverified AI currently cannot go. Competition policy and public infrastructure would keep markets open and ensure the Tool AI ecosystem doesn't depend entirely on a few large players. And hard limits are needed to provide physical constraints where soft incentives aren't sufficient.

The framework constrains the dangerous region of AI development. The vast majority of beneficial AI falls outside that region. Purpose-driven tools with bounded scope, meaningful human control, and appropriate assurance can be built, deployed, and improved with full support from this governance structure. The goal is a regulatory environment where building trustworthy tools is easier, cheaper, and more rewarding than building human replacements.

Six components work together to address the problems above.

1 - Autonomy and Responsibility

Those who maintain control bear appropriate responsibility for decisions made; those who cede control bear greater responsibility for the decision to cede it.

AI systems cannot bear legal responsibility. They have no assets, no liberty to forfeit, no conscience to deter. Responsibility must rest with the people and organizations who do have these things. The question is how to allocate it, and how to create incentives that favor controllable systems over uncontrolled ones.

The framework classifies AI systems into three categories based on the degree of human control actually exercised:

Tools operate under meaningful human direction. The human controller bears primary responsibility for outcomes. Developers are liable for defects or misrepresentation; deployers for providing adequate control infrastructure. Developers who build robust control mechanisms and accurately represent their systems' capabilities receive liability protection through safe harbors. The more controllable the tool, the clearer the protection.

Supervised Agents have operational autonomy but genuine human oversight: real-time monitoring, competence to evaluate, and authority to intervene. Liability is shared, weighted toward the supervisor. Crucially, "supervision theater" doesn't qualify: nominal oversight without genuine attention, competence, and intervention capacity is not meaningful supervision.

Autonomous Agents operate without meaningful human control, whether by design or because purported supervision is ineffective. Developers, deployers, and operators face joint-and-several liability, and strict liability for the highest-power systems. High-capability systems without certification are presumed to fall in this category.

The incentive structure is deliberate: everyone in the chain benefits from more controllable systems. Developers want certification for protection from liability. Deployers want operators who can exercise genuine oversight and take responsibility for systems under their genuine control. Operators want tools they can actually understand and direct. This liability structure also creates an incentive for developers and deployers to make tools rather than agents, in keeping with the Pro-Human Path preference for Tool AI. It is also constructed to match legal intuition and anticipate what courts would probably determine anyway – but without years of litigation and uncertainty.

Open-weight release does not exempt developers from responsibility. Models must be evaluated for dangerous capabilities before release, with reasonable measures to prevent misuse. Release of clearly dangerous models creates liability for foreseeable misuse. Open-weight thresholds are set below deployment thresholds because release is irreversible: once a model is in the wild, oversight is lost. Those who significantly fine-tune open-weight models become developers of the resulting system and inherit corresponding obligations.

Enhanced liability for children and vulnerable populations applies heightened standards: stricter duties of care, prohibition on manipulation techniques, special data protections, and a rebuttable presumption of harm where certain system behaviors affect vulnerable groups. Even small individual harms can aggregate to substantial liability when affecting large populations.

The framework also establishes personal criminal liability for executives who knowingly authorize development of prohibited systems or recklessly deploy systems without adequate safety evaluation. Corporate liability alone is insufficient when the potential harms are catastrophic.

AI systems that directly perform functions carrying fiduciary duties when performed by humans – an AI therapist, financial advisor, or legal advisor operating with significant autonomy rather than merely assisting a human professional – must meet those same fiduciary standards: duty of care, duty of loyalty, full disclosure of limitations and conflicts. And the responsibility accorded to fiduciaries does not go away: it is held by the AI developer and deployer. Where a human professional uses AI as a tool, the professional's existing fiduciary obligations apply to their use of that tool.

Enforcement operates through multiple channels. Regulatory bodies have investigation and penalty authority. But regulatory capacity is limited, so the framework also provides for private enforcement: individuals harmed by AI systems can sue developers, deployers, and operators; class actions address aggregate harms; prevailing plaintiffs recover attorney fees; and liability protections cannot be waived by contract or mandatory arbitration.

2 - Assurance Framework

A structured process for demonstrating that AI systems are safe, trustworthy, and under control, with different components serving different governance functions.

The current approach to AI safety is largely reactive: deploy first, discover problems through use, patch after the fact. We don't use this method for cars, planes, drugs, or almost any other product. The reason why is self-evident: it reacts to harms created, rather than prevents thems. The assurance framework provides the evidentiary infrastructure that the rest of the governance system relies on. It defines how claims about AI systems are substantiated, who evaluates them, and what the results are used for.

The framework applies across the full capability spectrum, but its requirements scale with risk. At the low end, developers self-certify with audit rights. As systems become more capable, requirements escalate through formal assurance cases reviewed by accredited certifiers, to enhanced scrutiny for systems elevated on multiple dimensions, to dual independent certification for the most capable systems. This is a gradient, not a binary threshold. It is designed to be nimble, proportional, and evolve quickly with the technology itself.

Four types of assurance case address different dimensions:

Safety and Security cases provide comprehensive risk assessment: identified harms, mitigation measures, and justification of residual risk. Above defined capability thresholds, these become hard legal requirements enforced through the compute governance framework below. A system that cannot demonstrate acceptable safety cannot be deployed.

Control cases demonstrate that meaningful human control exists, with specific mechanisms identified. These feed directly into the liability classification: a developer seeking safe harbor protection under the autonomy-responsibility framework must produce a credible control case. The strength of the control case determines whether a system qualifies as a Tool, Supervised Agent, or Autonomous Agent, with corresponding liability consequences. Developers at any capability level can benefit from producing a control case, since safe harbors reward it.

Trust cases validate that a system does what it claims to do, reliably, across its intended uses and under adversarial conditions. This includes also transparency and epistemic standards. Trust cases are part of the certification process at higher capability levels, but they also create direct market advantage at any level. Enterprise and government adoption of AI is blocked primarily by insufficient reliability, not insufficient capability. A strong trust case unlocks deployment in high-stakes settings where unverified systems cannot go.

Pro-Human evaluations assess whether a system serves human interests: whether it creates unhealthy dependency, employs manipulation, undermines user autonomy, or concentrates power. These are not legally mandated in the same way. They are published independently, carrying normative and reputational weight. Over time, market pressure and public expectation may make them effectively necessary for credible deployment.

Categorical exclusions apply regardless of assurance case quality. Some systems cannot be certified: prohibited architectures (recursive self-improvement, autonomous self-replication, resistance to shutdown) and intrinsically dangerous applications (autonomous weapons with lethal authority, mass surveillance infrastructure, social scoring systems, psychological manipulation systems, deceptive AI relationships). No safety case can make these acceptable.

Assurance cases are evaluated by accredited independent certifiers. The framework includes anti-capture mechanisms: revenue concentration limits prevent certifiers from depending on any single client, mandatory rotation requires periodic change of certifier, and certifier liability ensures accountability for negligent certification. This ecosystem does not yet exist and will need to be built, likely starting with pilot certifications in well-defined domains and expanding as standards and institutional capacity mature.

Placeholder

Intel's High Numerical Aperture Extreme Ultraviolet lithography tool is ready for calibration in its clean room. As software becomes increasingly difficult to monitor, the sheer physical scale and extreme scarcity of this hardware—which requires specialized facilities and immense energy—provide the 'natural lever' for global safety governance and hard limits on runaway capability scaling. (Image by Intel Foundry.)

3 - Hard Limits and Hardware Governance

Some risks require robust prevention, not just evaluation and incentives. Compute is the key enforcement lever.

The assurance framework manages risk through evaluation and certification. The autonomy-responsibility framework manages risk through liability incentives. But certain large-scale risks, particularly loss of control from systems that are simply too powerful, require harder constraints. Some actors will accept liability risk if the potential payoff is large enough. Some will evade evaluation if they can. Prevention, backed by real enforcement, is necessary for the tail risks that matter most.

For AI, compute is the natural lever. Unlike software, which copies freely, compute is physical. Chips are manufactured in a handful of facilities, owned by identifiable entities, located in specific jurisdictions, and require substantial energy to operate. The supply chain for advanced AI chips is extraordinarily concentrated. Moreover, cryptographic and on-chip security mechanisms can bind AI software to AI-specific hardware for additional leverage. This makes governance feasible in a way that governing software alone is not.

The framework establishes several mechanisms:

Prohibited development paths. Some paths are categorically prohibited regardless of compute level: development of superintelligence, recursive self-improvement without human authorization, autonomous self-replication, systems designed to resist shutdown or escape containment. These are not manageable; they must be foreclosed.

Model registration. All models above defined risk or capability thresholds must be registered, with training details, capability profiles, and deployment context disclosed. Registration creates accountability and enables oversight without prohibiting development.

Mandatory compute limits. Legal caps on training compute and inference compute rates, understood primarily as a backstop against runaway capability scaling. The primary delineation of which systems are viable comes through the safety and control framework; compute caps prevent the extreme tail. Training caps should not exempt "research" purposes: the primary risk is the system existing for any reason, not its deployment context.

Pre-deployment safety requirements. Systems above capability thresholds must pass safety and security evaluation per the assurance framework before deployment. Compute governance provides the enforcement lever; assurance cases provide the evaluation criteria.

Off-switch mandates. All systems above capability thresholds must have shutdown mechanisms independent of the system's cooperation, with graduated response levels and fail-safe defaults.

Open-weight release thresholds are set below deployment thresholds, reflecting the irreversibility of release. Once a model is publicly available, guardrails can be stripped, and there is no recall mechanism. Hard limits on what can be openly released are a necessary complement to deployment restrictions.

Compute thresholds are the starting point, not the endpoint. As algorithmic efficiency improves, the correlation between compute and capability will loosen. The framework anticipates an evolution: from compute-based limits (deployable now), to hardware-attested training (cryptographic proof of what code and data produced a model), to hardware-enforced capability constraints. Throughout this evolution, a capability backstop applies: systems demonstrating dangerous capabilities are prohibited regardless of compute used.

National security provisions allow AI systems to operate under separate classified oversight, but do not grant blanket exemptions. Prohibited architectures remain prohibited regardless of context. National security authority does not authorize mass surveillance of domestic populations. Where international agreements exist, national security programs operate within treaty constraints. The principle: national security requires protection from uncontrollable AI, not exemption from the rules preventing it.

4 - Privacy and Data Rights

Core values: cognitive liberty, economic fairness, power balance, and human dignity in an AI-intensive environment.

Privacy in the AI era serves values that go beyond traditional data protection. Cognitive liberty means the right to think, explore ideas, and form beliefs without (government or private) surveillance, or unauthorized inference into our mental states. Economic fairness means that people whose data, labor, and creative output fuel trillion-dollar AI systems should have meaningful control and a share of the value. Power balance means preventing the asymmetry where corporations know everything about individuals while remaining opaque themselves. Dignity means people are not data profiles to be optimized against.

The framework organizes protection in three layers:

Individual control includes data minimization (collect only what's necessary), genuine informed consent (not buried in terms of service), access and portability, deletion rights extending to training data, established requirement of affirmative consent for use of copyrighted material for training, and the right to contest automated decisions with meaningful human review.

Collective governance addresses the inadequacy of individual consent against concentrated corporate power. Data trusts and cooperatives enable communities to negotiate collectively. Public data commons serve research and public interest. Sectoral bargaining allows industries like healthcare or education to negotiate data terms as a group.

Categorical prohibitions recognize that some uses are impermissible regardless of consent: mass surveillance, real-time facial recognition in public spaces without specific justification, emotion recognition in coercive contexts, comprehensive social scoring, and exploitation of psychological vulnerabilities. Government surveillance is also constrained: AI analysis of communications, movements, or behavior patterns requires individualized warrants, and governments cannot bypass warrant requirements by purchasing equivalent analysis from private parties.

AI-specific provisions extend these protections to new categories: data generated through inference about individuals, data derived from personal information including embeddings and profiles, and personal data encoded in model weights. Synthetic data derived from personal data remains subject to the original data rights. Core data rights cannot be waived by contract.

The good news is that many of the violations of people's privacy rights are due to policy and not technology. As discussed under technologies, there are well-established mechanisms of differential privacy and other cryptographic methods – including some AI-based ones – that can circumvent many tradeoffs (in privacy versus security for example) that are used as excuses for over-collection of data. And the over-collection that has no excuse simply should not be allowed.

5 - Competition and Market Structure

Core concern: without countermeasures, AI markets will consolidate into monopolies that become too powerful to check.

The Tool AI path helps with competition. Purpose-built systems serving specific needs are more amenable to diverse, competitive markets than a winner-take-all race for one general-purpose system. Calling off this race would dramatically curtail the drive toward power concentration. But large-platform dynamics create strong centralizing pressures even for tools: network effects, data accumulation, compute concentration, vertical integration. Left alone, these forces produce the same kind of unchecked private power the framework aims to prevent.

Structural constraints prevent dominant entities from controlling multiple chokepoints in the AI stack. A company that dominates compute infrastructure should not also control the leading AI services. A company providing foundation models should not be able to foreclose competing applications built on those models. Mergers that substantially increase concentration should face heightened scrutiny and presumptive prohibition above defined thresholds.

Behavioral constraints ensure that dominant platforms enable rather than suppress competition. Interoperability mandates, documented APIs on reasonable terms, open standards for portability, and non-discrimination requirements prevent incumbents from locking out competitors. Users can move their data between providers.

Algorithmic accountability requires dominant platforms to disclose what their recommendation systems actually optimize for, submit to independent audits verifying that actual behavior matches disclosed targets, and provide accessible alternatives to algorithmic curation. For systems accessible to minors, engagement optimization, behavioral profiling, and amplification of content known to harm the mental health of young people are prohibited. Platforms must integrate with public provenance infrastructure, displaying indicators of content origin and epistemic context.

Access guarantees keep the field open for new entrants: tiered compliance that doesn't crush small developers, public AI compute resources and datasets, and support for open-weight models within appropriate safety constraints.

Placeholder

United States Trade Representative and Secretary of the Treasury, left, meet with Chinese Vice Premier in April 2019, during trade discussions at the USTR offices in Washington, D.C. (Official White House Photo by Andrea Hanks)

6 - International Coordination

Core logic: The AI race serves companies more than countries. Verifiable mutual restraint is in everyone's interest, including the two countries leading the race.

Domestic governance alone is insufficient. Without international coordination, companies can "jurisdiction-shop" to avoid constraints, and geopolitical rivalry and competition drive both sides toward increasingly dangerous development. Desire to avoid large-scale AI risks can drive international coordination on some safety measures. But the strategic logic for coordination is stronger than this because unlike other powerful technologies advanced AI can grant power – but it can also take it away.

Because systems of unprecedented capability are being developed in private hands, as described in Dystopian Dynamics, it may threaten governments directly. The chain runs: companies escape government control, then companies gain power over governments, finally power is handed to AI, or taken by it. At the end, no human institution is in charge.

This threat structure is symmetric. The US government faces it from US companies; the Chinese government faces it from Chinese companies. Neither benefits from the other's companies succeeding in this trajectory. Both have self-interested reasons to constrain AI development that have nothing to do with altruism.

Hardware provides the common lever. Both AI superpowers produce the advanced chips on which frontier AI depends. Both have incentives to ensure those chips are not used in ways that threaten their own security. This creates a shared interest in harmonized hardware governance, similar to how the US and Soviet Union agreed not to undercut each other on oversight of nuclear materials.

A plausible trajectory runs through five steps: independent recognition by both leading AI powers that the race is undesirable; domestic hardware governance with remote verification; confidence-building measures (hotlines, public information sharing, joint research on verification); a bilateral agreement on baseline requirements; and eventually a multilateral treaty. These steps need not be strictly sequential. Countries can begin building verification mechanisms, forming coalitions, and developing model frameworks before major power agreement materializes. And confidence-building can start right away.

Redlines suitable for codification focus on behaviors posing unacceptable risk regardless of intent: design or use of weapons of mass destruction, autonomous recursive self-improvement, autonomous self-replication, evaluation gaming and alignment faking, loss of human control, and highly effective manipulation.

How the Components Work Together

These six components are designed as an interlocking system, not a menu of independent options.

The autonomy-responsibility framework creates market incentives favoring controllable tools, but those incentives only work if there is a credible way to demonstrate control. The assurance framework provides that mechanism: control cases feed into liability classification, safety cases are enforced as hard requirements above capability thresholds, trust cases unlock market access, and pro-human evaluations build normative expectations. Assurance applies across the capability spectrum with escalating rigor, so the incentive structure works at every level, not just for frontier systems.

Hard limits provide physical constraints that incentives alone cannot. Some actors will accept liability exposure if the potential payoff is large enough. Compute caps and prohibited architectures create boundaries that don't depend on anyone's cost-benefit calculation. The assurance framework provides the evaluation criteria that hard limits enforce.

Privacy and data rights constrain manipulation and exploitation at the source, complementing assurance requirements that evaluate systems for these harms before deployment. Competition policy ensures that governance doesn't inadvertently entrench incumbents by creating compliance burdens only large players can bear, while public infrastructure and access guarantees actively support the Tool AI ecosystem.

International coordination extends all of these across borders, preventing the domestic framework from being undermined by jurisdiction- shopping or geopolitical race dynamics. Hardware governance provides the shared enforcement lever.

Remove any component and gaps emerge. Liability without assurance is unverifiable. Assurance without hard limits can't prevent the most dangerous systems. Domestic rules without international coordination get arbitraged. The framework is designed to hold together.

Placeholder

The EU AI Act is adopted by the European Union. (Image by European Commission.)

How We Get There

A governance framework this comprehensive might seem unrealistic given the current political landscape. We are not proposing that all of this can be enacted tomorrow. We are proposing that this is what is actually needed, and that the conditions for enacting it are developing faster than many assume.

The political window is opening. Public concern about AI harms is growing. Governments are beginning to recognize the threat that unchecked AI companies pose to state authority. The insurance industry faces mounting pressure to price AI risk. Labor movements are mobilizing. AI incidents will accelerate these trends. When the moment arrives for serious governance, the response should be guided by a coherent framework rather than assembled ad hoc.

Chokepoints exist. The resources required for frontier AI development are physical, concentrated, and governable. Advanced chips are manufactured in a small number of facilities. Training runs require enormous energy with physical infrastructure. The talent pool is identifiable. Each is a point where governance can gain traction. This stands in contrast to technologies like software or information, where control points are few and copying is free.

Precedent exists. Although AI is different in key ways, it is also a technology and a set of products, and there is plenty of precedent for governing the creation of new technologies and products, from toys up to jumbo jets. And powerful but dangerous technologies have been meaningfully constrained before. Nine states have nuclear weapons, not ninety. The Biological Weapons Convention is imperfect, but the norm is real and most capable states have not developed bioweapons. Reproductive cloning and human germline engineering are technically feasible – and could be very profitable – but are effectively prohibited. These regimes demonstrate that imperfect constraint is vastly better than no constraint.

The components can be phased. Some elements are enactable now: bot-or-not disclosure, child protection provisions, basic liability clarification, compute reporting requirements. Others build on those foundations: the assurance ecosystem, starting with pilot certifications in well-defined domains and expanding as standards and institutional capacity develop. Hardware governance and international coordination develop in parallel, starting with confidence-building measures and moving toward formal agreements as verification mechanisms mature.

Reinforcing dynamics help. Creating viable Tool AI alternatives makes the case for governance stronger. Governance creates market conditions favoring Tool AI. Closing off dangerous paths makes the Pro-Human Path more attractive, and the Pro-Human Path's development helps relieve the pressure to push (or allow) for development in dangerous directions. Each step makes subsequent steps easier.

The window matters. The further AI development proceeds along the current trajectory, the harder redirection becomes. Companies accumulate political power. Systems become entrenched in infrastructure. Norms solidify. Putting the right framework in place while the technology and the industry are still taking shape is substantially easier than retrofitting it later.

The goal of this governance framework is not to slow AI down.

It is to steer AI development toward systems that are controllable, verifiable, and pro-human, while foreclosing paths that concentrate power, threaten general safety and wellbeing, or escape human oversight. The rules create the conditions under which the Pro-Human Path can win.

Share this page: