StableMindEnterprise Evaluation
Sign inOpen evaluation room

SIO27 · Enterprise Evaluation Room

Evaluate the boundary.
Before you approve the deployment.

Enterprise evaluation of agentic AI should not begin with a feature checklist. It should begin with the action the organization is unwilling to delegate today, then ask where power would come from, how far it may travel, which consequences require capacity, where credentials exist, how execution can be interrupted, and what evidence would be strong enough to support a decision after the action.

01 · Evaluation framework

Six disciplines. One question: what has to be true before this machine should be allowed to act?

StableMind separates evaluation concerns because combining them too early makes diligence look greener than reality. Security can pass while authority design is incomplete. Architecture can be coherent while deployment locality is wrong. A procurement team can approve commercial terms without approving consequential execution. Evidence can prove a synthetic evaluation without proving a customer outcome. The Enterprise Evaluation Room keeps those facts adjacent without letting one substitute for another.

01

Security

Inspect identity separation, credential custody, revocation, tenant isolation and whether any convenience layer can bypass the authority boundary.

02

Architecture

Trace delegated authority, deterministic ActionGate decisions, Action Permits, Execution Fabric, Guardian, Evidence and the boundaries between them.

03

Deployment

Evaluate locality, customer control, air-gap requirements, read-only shadow observation, bootstrap verification, rollback and teardown.

04

Procurement

Separate contractual/commercial acceptance from technical evidence, security assurance, machine authority, production approval and realized value.

05

Evidence

Define what must be precommitted, observed independently, retained, attributed and strong enough to support Proof of Consequence.

06

Pilot Design

Anchor the pilot to a refused workflow, establish a counterfactual baseline and define customer-owned graduation criteria before writes exist.

02 · Security & authority

A security review should prove more than “the agent authenticated.”

Traditional enterprise controls answer valuable questions: who authenticated, what token was issued, what network path was used, which secrets are vaulted, and what service account can reach the API. Machine Authority Infrastructure adds a different question: where did the right to exercise this exact consequential power come from?

For StableMind, a commercial account, organization role, subscription, workspace membership, API client, connector, SDK, secret, credential, or reachable endpoint is not machine authority. The evaluation should verify that effective power originates in explicit delegated authority, remains bounded by scope and time, attenuates through descendants, can be revoked independently, and is evaluated against the exact semantic action before technical capability is materialized.

A buyer should also inspect failure behavior. If an ancestor delegation is revoked, descendants should lose current validity. If required evidence is absent, Proof of Consequence should fail to UNVERIFIABLE rather than becoming a optimistic success. If a customer deployment requires local credential custody, a managed user interface should not quietly pull those raw credentials into a remote plane. If the requesting agent ignores an instruction to stop, Guardian must not depend on that agent’s cooperation.

03 · Architecture & deployment

Trace the request all the way from mandate to consequence, then inspect every handoff.

A meaningful architecture review should follow the governing chain rather than review each service in isolation: delegated authority → semantic action request → deterministic decision → exact permit → consequence reserve where required → JIT credential → certified execution → receipt → independent evidence → Proof of Consequence. StableMind intentionally keeps responsibilities separate so an implementation cannot obtain a favorable conclusion merely because one component says another succeeded.

01Origin

Who grants authority, to whom, for which action/resource/destination, under what limits, validity period and inherited constraints?

02Decision

Does ActionGate evaluate the current semantic request deterministically against current authority and policy rather than model confidence?

03Capability

Does technical capability appear only after a bounded permit, with JIT credentials and a certified executor unable to broaden the permit?

04Interruption

Can Guardian narrow, suspend, revoke, quarantine or coordinate recovery independently, without creating new power?

05Evidence

Are executor receipts separated from independent postcondition observations and final Proof-of-Consequence state?

Deployment review then asks where each plane lives. StableMind’s sealed engineering lineage includes customer-managed Kubernetes, private cloud, air-gapped enclave, managed service and hybrid customer-local enforcement profiles. A profile being marked production-eligible in source is an engineering property, not proof that a customer has deployed it. The evaluator should choose the profile that fits locality, isolation, credential custody, observability and evidence-retention requirements, then verify those requirements in the customer environment.

04 · Procurement & Public Truth

Do not let a diligence packet become an evidence laundering machine.

Enterprise procurement often compresses several decisions into one word: approved. StableMind deliberately refuses that compression. A legal agreement is not a security approval. A security questionnaire is not proof that a deployment is operating correctly. A successful synthetic ActionGate ceremony is not a live customer action. A sealed production-deployment build is not customer production evidence. A Proof of Consequence is not automatically customer acceptance, settlement, revenue recognition or realized value.

The Enterprise Evaluation Room therefore carries evidence class with the evaluation artifact. Architecture is labeled architecture. Source-backed engineering is labeled source-backed engineering. Synthetic demonstrations are labeled synthetic. Historical build evidence remains build-scoped. External production proof must come from attributable external evidence and cannot be manufactured by completing a checklist on stablemind.io.

This same discipline protects the buyer. Internal review teams can record their own conclusions and exceptions without StableMind converting those conclusions into marketing claims. The workspace room can become READY_FOR_INTERNAL_REVIEW when the six evaluation disciplines have enough customer-defined inputs, but that state has no procurement, deployment, security-certification or production-readiness authority.

05 · Evidence & pilot design

Start with the workflow the customer refuses to automate, then make graduation falsifiable.

The strongest evaluation wedge is not “show us everything StableMind can do.” It is one consequential workflow the organization currently refuses to delegate. Define the exact semantic action and consequential target. Identify the candidate authority source and human approval boundary. Select the relevant consequence dimensions. Precommit the expected postcondition and prohibited collateral effect. Then observe the existing workflow in customer-controlled read-only shadow mode and compare what ActionGate would have done.

A

Challenge

SIO22 creates a non-authorizing Challenge Manifest. Missing target, authority source, approval boundary, consequence exposure or proof requirement remains incomplete rather than guessed.

B

Shadow

SIO23 can project WOULD_PERMIT, WOULD_NARROW, WOULD_REQUIRE_HUMAN_AUTHORITY, WOULD_DENY or AUTHORITY_UNRESOLVED from minimized read-only observation. Counterfactual value is not realized value.

C

Graduation

The customer defines the evidence, operational, authority, deployment and connector conditions required before any later move toward consequential execution. SIO27 cannot grant graduation.

Pilot success criteria should therefore be concrete enough to fail. Examples include: required authority lineage was unresolved in fewer than an agreed fraction of eligible observations; changed consequential targets reliably triggered HUMAN_AUTHORITY; requests outside delegated limits were narrowed or denied; no write scope existed during shadow; required verifier observations were attributable; and all exceptions were explainable to a human reviewer. Those are examples of evaluation design, not claims that any customer has already achieved them.

05A · Make the review operational

Define who can say yes, who can say no, and what evidence survives the meeting.

A serious enterprise evaluation is not complete when the product team understands the architecture. It is complete enough for internal review only when the organization knows which functions own which questions, which evidence each function expects, how exceptions are recorded, and which conclusion belongs to which decision. Security may own credential custody and isolation. Architecture may own delegation and enforcement boundaries. Procurement may own commercial terms and supplier risk. Legal or compliance may own jurisdictional obligations. Operations may own rollback and incident response. Finance may own payment or reserve controls. None of those owners should be silently replaced by a single StableMind score.

For that reason, SIO27 does not publish a composite “enterprise readiness” percentage. A percentage would invite false equivalence between unlike facts. A deployment-locality exception cannot be canceled out by excellent SDK documentation. Missing independent verifier evidence cannot be offset by a strong security questionnaire. A procurement approval cannot compensate for an unresolved authority lineage. Each discipline should retain its own open questions, evidence references, exceptions, owners and decision rights until the customer chooses how to govern them.

The same rule applies to diligence artifacts. A packet should be reproducible enough that a reviewer joining later can see which StableMind source artifact supported a statement, which claim was architectural, which demonstration was synthetic, which control was tested against the current build, and which conclusion was supplied by the customer. That makes the evaluation more durable and also prevents the company website from becoming an accidental source of invented customer proof.

When the review moves toward a pilot, the organization should define the transition criteria before connecting consequential capability: required observation period, acceptable unresolved-authority rate, expected human-escalation behavior, prohibited write scopes during shadow, evidence retention, incident response, connector certification expectations, deployment locality and explicit customer authorization for any later execution stage. SIO27 can organize those questions. It cannot answer them on the customer’s behalf.

06 · Eighteen source-backed evaluation questions

Use the room as a diligence map, not a checkbox score.

EVAL-01
Authority boundary

Can commercial identity, API access, or credentials create machine authority?

SECURITY
EVAL-02
Credential custody

Where do raw consequential credentials exist?

SECURITY
EVAL-03
Revocation

Can machine power be interrupted independently of the requesting agent?

SECURITY
EVAL-04
Decision independence

Is authorization deterministic and separate from model inference?

ARCHITECTURE
EVAL-05
Delegation lineage

Can evaluators trace where power came from?

ARCHITECTURE
EVAL-06
Execution separation

Can the governance client itself execute the action?

ARCHITECTURE
EVAL-07
Locality

Can enforcement run where the customer requires it?

DEPLOYMENT
EVAL-08
Shadow first

Can the refused workflow be observed before write capability exists?

DEPLOYMENT
EVAL-09
Rollback and teardown

Can the evaluation boundary be removed cleanly?

DEPLOYMENT
EVAL-10
Claims discipline

Which claims are architecture, source-backed engineering, synthetic evidence or external proof?

PROCUREMENT
EVAL-11
Commercial boundary

Do subscription, ownership or workspace roles create authority?

PROCUREMENT
EVAL-12
Evaluation exit criteria

What does a completed evaluation actually approve?

PROCUREMENT
EVAL-13
Receipt versus proof

Does executor success independently prove consequence?

EVIDENCE
EVAL-14
Missing evidence

What happens when required verifier evidence is absent?

EVIDENCE
EVAL-15
Evidence custody

Who holds evaluation artifacts and how are they attributable?

EVIDENCE
EVAL-16
Refused workflow

Is the pilot anchored to an action the customer currently refuses to automate?

PILOT DESIGN
EVAL-17
Counterfactual baseline

Can the customer measure what ActionGate would have changed without executing?

PILOT DESIGN
EVAL-18
Graduation gate

What must be true before consequential execution is considered?

PILOT DESIGN

Machine-readable version: /enterprise-evaluation/requirements.json. The questions describe what an evaluator should inspect; they are not a public certification scheme and do not confer conformance, customer approval or authority.

07 · Evaluation handoff

Build a customer-owned dossier without turning the dossier into a decision.

The authenticated Evaluation Room lets an organization choose its evaluation disciplines, define internal reviewers, record evidence expectations and export a browser-local dossier for internal review. Customer-entered content stays in the browser session in SIO27. No procurement system, ticketing system, identity provider or customer environment is connected by completing the room.

A dossier reaching READY_FOR_INTERNAL_REVIEW means only that the selected evaluation dimensions are sufficiently defined to hand to the customer’s own review process. It does not mean StableMind has approved the customer, the customer has approved StableMind, or a consequential AI agent is ready for production.