Systems Atlas/ AI, Work & Organizations

State of Knowledge

AI, Work & Organizations

An integrated current snapshot, not a changelog and not a claim of completion.

Version
v2.1
State date
2026-09-27
Research stage
PRIVATE OPERATIONAL DISCOVERY

Current thesis

AI changes the feasible set and relative cost of organizational action. Realized effects depend on the resulting configuration of actors, incentives, information, authority, task boundaries, workflows, capability stocks, and external constraints. That configuration produces outcomes and distributes value and burdens; those consequences alter future behavior, work allocation, capabilities, and organizational structure. A causal arrow may be supported by credible causal identification. A workflow-specific binding constraint is treated as intervention-validated only when it is specified prospectively, directly relaxed, passes a mechanism check, and produces the predicted downstream response.

  1. AI changes the feasible set
  2. configuration changes
  3. constraints and outcomes emerge
  4. the system learns and reconfigures

Knowledge frontier

What current evidence supports ↔ what could still change the map

Currently supported

Task-level AI gains do not guarantee end-to-end workflow or organizational gains.

SUPPORTEDcentral
CLM-02Benchmark capability does not determine organizational productivity.
SUPPORTEDACTIVEcentral

Limit: Configuration-specific effect sizes remain heterogeneous.

Investigation path

  • INV-01 · Capability, Tasks & Organizational Consequence
  • INV-02 · Workflow Propagation

Selected evidence

  • E02 · BCG / HBS jagged frontierrandomized field experiment

    Strong speed and quality gains appeared inside the model's task frontier, while correctness fell on an outside-frontier task.

    Does not establish: General organizational productivity.

  • E05 · METR experienced OSS developer RCTrandomized controlled trial

    Participants using early-2025 AI tools were about 19% slower while believing they were faster.

    Does not establish: A universal effect of coding assistants.

  • E06 · Integrated knowledge-work AIrandomized workplace field experiment

    Active users spent about two fewer hours per week on email and less time outside normal hours, without detectable broad work-composition change.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-02

CLM-05Local task gains do not necessarily propagate into workflow improvement.
SUPPORTEDACTIVEcentral

Limit: Propagation magnitude and dominant mechanisms vary by workflow.

Investigation path

  • INV-02 · Workflow Propagation

Selected evidence

  • E06 · Integrated knowledge-work AIrandomized workplace field experiment

    Active users spent about two fewer hours per week on email and less time outside normal hours, without detectable broad work-composition change.

  • E07 · Startup AI opportunity mappingrandomized field experiment

    Broader mapping of AI into production increased discovered use cases, completed tasks, paying-customer probability, and upper-tail revenue.

    Does not establish: The same effect in established enterprises.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-05

What would still change the map

Which configuration variables prospectively predict when local gains propagate?

Next discriminating evidence: Repeated interventions that measure the workflow graph, state, authority, actor incentives, and true downstream outcome.

Open unknowns: U13

Currently supported

State quality can be produced by behavior, incentives, authority, or control rather than only by technical capture limits.

CONTEXT-CONDITIONALstate formation
CLM-38Some constraints are endogenously produced by organizational configuration.
SUPPORTEDACTIVEmechanism

Limit: Prevalence across workflows is unknown.

Investigation path

  • INV-17 · Incentives, Information & Constraint Formation

Selected evidence

  • E65 · AI reliance under evaluator observabilityrandomized worker experiment

    Making AI reliance observable to an evaluator changed workers' use of the same AI recommendations and reduced performance.

    Does not establish: That incentive/evaluation effects dominate enterprise deployments.

  • E66 · Authority and soft-information productionorganizational empirical study in financial services

    Greater authority for loan officers changed effort devoted to producing and using soft information.

    Does not establish: The magnitude of this mechanism in AI-enabled workflows.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-38

CLM-43State quality can be an endogenous behavioral outcome rather than only an information-system property.
CONTEXT-CONDITIONALACTIVEstate formation

Limit: Field prevalence in AI-enabled workflows is unknown.

Investigation path

  • INV-17 · Incentives, Information & Constraint Formation

Selected evidence

  • E65 · AI reliance under evaluator observabilityrandomized worker experiment

    Making AI reliance observable to an evaluator changed workers' use of the same AI recommendations and reduced performance.

    Does not establish: That incentive/evaluation effects dominate enterprise deployments.

  • E66 · Authority and soft-information productionorganizational empirical study in financial services

    Greater authority for loan officers changed effort devoted to producing and using soft information.

    Does not establish: The magnitude of this mechanism in AI-enabled workflows.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-43

What would still change the map

How often are apparent state failures incentive-produced in recurring AI-enabled workflows?

Next discriminating evidence: FIELD-01-style information × incentive discrimination with successful mechanism checks.

Open unknowns: U01, U33

Currently supported

AI can change task economics and, in bounded settings, transform dependencies that were previously supplied by human teaming or specialization.

SUPPORTEDbounded workflow design
CLM-37Workflow configuration can be endogenous to AI capability and task economics.
SUPPORTEDACTIVEbounded

Limit: General magnitude remains unknown.

Investigation path

  • INV-18 · Workflow Endogeneity & Task Rebundling

Selected evidence

  • E68 · P&G professional-team AI experimentpreregistered randomized field experiment

    AI-assisted individuals achieved performance comparable to conventional two-person teams on studied innovation tasks and reduced some functional-specialization gaps.

    Does not establish: That operational teams generally become unnecessary.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-37

What would still change the map

How much long-run value comes from graph recomposition rather than faster nodes?

Next discriminating evidence: NO AI + OLD WORKFLOW / AI + OLD WORKFLOW / NO AI + REDESIGNED / AI + REDESIGNED factorial.

Open unknowns: U04

Currently supported

Assisted performance and durable human capability are distinct, and interaction design can change learning effects.

SUPPORTEDcapability reproduction
CLM-08Assisted performance and durable human capability are distinct.
SUPPORTEDACTIVEcentral

Limit: Long-run workplace effects remain undermeasured.

Investigation path

  • INV-04 · Human Expertise, Apprenticeship & Oversight Capacity
  • INV-19 · Capability Stocks & Expertise Reproduction

Selected evidence

  • E11 · Anthropic coding-skills RCTrandomized learning experiment

    AI-assisted participants had lower subsequent mastery; independent debugging and explanation-oriented use preserved more learning.

    Does not establish: Broad workplace deskilling.

  • E12 · Patent-law RCTpreregistered workplace experiment

    AI improved drafting; durable unaided judgment gains were concentrated among senior lawyers.

  • E72 · Unrestricted AI assistance versus guarded tutoringfield experiment in education

    Unrestricted assistance improved assisted practice but worsened later unaided performance relative to control, while a tutor design preserving more cognitive work largely avoided the penalty.

    Does not establish: Long-run workplace deskilling.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-08

What would still change the map

What happens to diagnosis, verification, and recovery after sustained AI-mediated work?

Next discriminating evidence: Longitudinal capability-stock testing under normal AI, no AI, misleading AI, unfamiliar cases, and failure conditions.

Open unknowns: U08, U09, U10

Currently supported

Configuration is a plausible causal organizing object that preserves mechanisms hidden by task-only analysis.

INFERREDcentral
CLM-44Configuration is currently a more useful causal organizing object than constraint alone.
INFERREDACTIVEcentral

Limit: The configuration layer must earn its complexity through prospective prediction and intervention value.

Investigation path

  • INV-17 · Incentives, Information & Constraint Formation
  • INV-18 · Workflow Endogeneity & Task Rebundling
  • INV-19 · Capability Stocks & Expertise Reproduction

Selected evidence

  • E65 · AI reliance under evaluator observabilityrandomized worker experiment

    Making AI reliance observable to an evaluator changed workers' use of the same AI recommendations and reduced performance.

    Does not establish: That incentive/evaluation effects dominate enterprise deployments.

  • E68 · P&G professional-team AI experimentpreregistered randomized field experiment

    AI-assisted individuals achieved performance comparable to conventional two-person teams on studied innovation tasks and reduced some functional-specialization gaps.

    Does not establish: That operational teams generally become unnecessary.

  • E72 · Unrestricted AI assistance versus guarded tutoringfield experiment in education

    Unrestricted assistance improved assisted practice but worsened later unaided performance relative to control, while a tutor design preserving more cognitive work largely avoided the penalty.

    Does not establish: Long-run workplace deskilling.

Source: systems/ai-work-organizations/CLAIMS.md · CLM-44

What would still change the map

Does configuration actually improve out-of-sample prediction and intervention design?

Next discriminating evidence: Prospective organizational interventions compared against a task-capability baseline.

Open unknowns: U12, U13

Currently supported

Imported mechanisms are not automatically portable across countries, firms, or operating substrates.

SUPPORTEDgeography / institutions

What would still change the map

Which mechanisms transfer to Malaysia and Southeast Asia, especially across formal enterprise and owner-operated MSME configurations?

Next discriminating evidence: Local field cases that measure state production, incentives, authority/control, task boundaries, and accepted outcome.

Open unknowns: U14

What changed

v2.1 changed causal ordering, not merely vocabulary

What we no longer claim

Revision is part of the evidence record

CLM-32REJECTED

Workflow is fixed before AI propagation is evaluated.

Current replacement: Workflow configuration can be endogenous to AI capability, task economics, task boundaries, and managerial choice.

CLM-33REJECTED

Expertise can be represented as one scalar capability stock.

Current replacement: Use distinct capability stocks such as production, diagnosis, verification, recovery, transfer, calibration, explanation, and teaching.

CLM-35REJECTED

Poor organizational state primarily indicates a technical information-architecture failure.

Current replacement: Test technical capture, incentives, distribution, strategic withholding, authority, and control as rival origins.

CLM-36REJECTED

Authority is adequately represented as one undifferentiated variable.

Current replacement: Separate formal authority, practical control, informational control, accountability burden, economic upside, and discretion value where material.

CLM-31NARROWED

A new problem appearing after another constraint is relaxed demonstrates bottleneck migration.

Current replacement: Predict the next constraint before intervention, verify the intended mechanism changed, observe its signature, then test it directly.

CLM-34NARROWED

Value distribution is purely downstream of production.

Current replacement: Expected value and burden distribution can change actor behavior before value is realized.

Major unresolved questions

Unknowns are ranked by their ability to change the representation

  1. U01

    How often are apparent state or information problems primarily produced by actor incentives, burden, accountability, or control?

    It determines whether technical data fixes address root causes or symptoms.

    Next evidence: Recurring workflow studies that manipulate information and actor incentives separately.

  2. U04

    How strongly do economic task boundaries change with AI capability, reliability, and cost?

    It determines whether AI value comes mostly from faster nodes or from changing the workflow graph.

    Next evidence: AI × workflow-redesign factorials that separate node acceleration from rebundling.

  3. U08

    How does diagnostic capability change after sustained AI-mediated work?

    Normal-path productivity can hide declining ability to understand novel failures.

    Next evidence: Longitudinal tests under normal AI, no AI, misleading AI, and novel cases.

  4. U09

    How does recovery capability change when humans handle fewer normal and failed cases directly?

    Deep automation can improve normal operation while weakening emergency recovery.

    Next evidence: Longitudinal recovery scenarios and controlled failure exercises.

  5. U10

    Can strong verification be maintained after direct production capability declines?

    Many proposed oversight architectures assume verification remains reliable even when execution practice disappears.

    Next evidence: Misleading-AI tests after sustained changes in production experience.

  6. U12

    Can next constraints be specified prospectively and predicted better than a baseline or null model?

    If not, constraint-trajectory language adds little scientific value.

    Next evidence: Prospective constraint-prediction benchmarks against explicit baselines.

  7. U13

    Do configuration variables materially improve out-of-sample prediction and intervention design beyond task-level AI capability?

    This is the strongest empirical test of whether v2.1 earns its added architecture.

    Next evidence: Repeated prospective organizational interventions across different workflow configurations.

  8. U14

    Which causal mechanisms transfer from U.S. or European evidence to Malaysia and Southeast Asia, and which do not?

    Wages, firm size, language, informality, management, and digital depth can change the economics and mechanism.

    Next evidence: Local field studies rather than direct portability assumptions.

  9. U33

    Can one eligible exception case be joined to the actual accepted downstream organizational state?

    FIELD-01 cannot distinguish process completion from true outcome without case-level linkage.

    Next evidence: Private operational evidence from an eligible host-workflow configuration.

  10. U34

    How many genuinely independent workers, teams, sites, queues, or accounts handle eligible cases?

    High transaction volume does not guarantee enough independent experimental units.

    Next evidence: Private worker/team/assignment structure from an eligible host-workflow configuration.

Strongest falsifier

The configuration layer has to earn its complexity

v2.1 should be simplified if repeated prospective organizational studies show that task-level AI capability predicts organizational outcomes almost as well as actor incentives, task boundaries, workflow configuration, state endogeneity, and capability stocks.

What we are testing next

FIELD-01 is the first direct causal-discrimination program

It asks whether a recurring exception workflow stalls because of information, incentives, their interaction, or neither, while AI capability is held fixed. It is still pre-operational.

Inspect FIELD-01 →

Previous states

v1 → v2 → v2.1

v1RETIRED

AI was represented as uneven capacity expansion inside a gated adaptive work system; bottleneck migration was a central explanatory idea.

v2RETIRED

The map became explicitly recursive: capacity changes propagated through state, evidence, authority, coordination, execution, economics, labor, and physical reality, then altered the next system state.

v2.1CURRENT

The causal center moved upstream to organizational configuration. Constraints, workflow, state production, capability stocks, and value/burden distribution can be endogenous.

v1 → v2system dynamicsADDED

Before

A gated work system with bottleneck migration as the main dynamic.

After

An explicitly recursive system with state, authority, economics, labor, physical reality, feedback, stocks, delays, and operating regimes.

Why: Investigations 02–16 showed that local productivity, organizational state, authority, learning, economics, resilience, and physical execution could not be treated as one linear propagation chain.

Consequence: The map became a coupled adaptive system in which earlier AI effects changed the conditions for later AI effects.

Trigger investigations: INV-02, INV-05, INV-06, INV-08, INV-10, INV-12, INV-14, INV-15, INV-16

v2 → v2.1organizational configurationMOVED

Before

Capacity passed through largely pre-existing propagation constraints.

After

Feasible actions alter organizational configuration, and that configuration can produce or expose constraints.

Why: The v2 model was too flexible retrospectively and placed several variables in conflicting causal roles.

Consequence: Configuration becomes the main causal organizing object and must earn its complexity prospectively.

Trigger investigations: INV-17, INV-18, INV-19

Inspect the full semantic diff →