Governance and Security (G)

This page corresponds to §9 of Agent Harness Engineering: A Survey. Governance and Security is the second layer the taxonomy promotes to first-class status. It asks how agent behavior is constrained, made safer, and made accountable.

Agents now execute shell commands, send email, submit code, browse sites, and call third-party APIs. Production systems therefore need answers to two questions: under what constraints may the agent act, and who is accountable when those constraints fail?

Permission Models And Identity

Access control for agents is harder than traditional RBAC/ABAC because the needed tool set often depends on a natural language task that was not known at deployment time.

GranularityMechanismExamples
Static permission boundaryFixed allow/deny lists and sandbox limitsCodex restricted sandboxes, Gemini CLI allow/deny
Context-dependent privilege controlEvaluate predicates over tool, arguments, environment, and task contextProgent, Conseca
Agent identity and delegationBind user intent to agent authority and audit chainAuthenticated Delegation, SAGA, IsolateGPT
Credential managementKeep secrets in a vault; expose placeholders to the modelSkyvern
Web-level permission coordinationSite-declared action permissions and rate/confirmation requirementsagent-permissions.json

Open problems include portable policy languages, task-scoped identities, and credential lifecycles for long-running sessions.

Lifecycle Hooks

Permission models define what is allowed. Lifecycle hooks define when checks fire.

HookLocationPurposeExamples
H1 input guardrailBefore the LLMDetect injection payloads in user input or retrieved contentPromptShield, DataSentinel
H2 action validationBefore tool executionCheck proposed actions and multi-agent control-flow transitionsShieldAgent, ControlValve
H3 information flow controlAfter tool executionTrack provenance and prevent untrusted data from influencing control flowCaMeL
H4 human-in-the-loopBefore consequential actions or at handoff pointsGate destructive or out-of-scope actions on approvalCodex, Gemini CLI, Cursor, OpenHands

Human approval is not free. Prior permission-UI research shows that users often ignore or misunderstand dialogs; agent approval prompts face similar habituation risks. Frequent prompts teach reflexive approval, while sparse prompts leave coverage gaps.

Hook APIs are heterogeneous, and stacked hooks can interact badly. An upstream sanitizer might remove the signal a downstream detector depends on.

Component Hardening

Hooks enforce policy around the loop. Component hardening makes the model and tools less vulnerable before hooks fire.

  • Model hardening: instruction hierarchy and SecAlign train models to prioritize privileged instructions over untrusted data.
  • Classifier runtime hardening: Llama Guard-style models screen input/output with configurable taxonomies.
  • Tool and MCP security: ETDI signs and versions tool definitions; SAFEFLOW adds protocol-level information flow control and transactional rollback.
  • Supply chain hardening: package hallucination and slopsquatting show that agents can introduce compromised dependencies even when the immediate tool call looks valid.

No single hardening layer covers all threats. Hardened models can still misuse compromised tools; signed tools do not prevent a jailbroken model from choosing unsafe actions.

Declarative Constitutions

Governance rules become more inspectable when externalized:

LevelMechanismSignal
Training-time constitutionShape model alignment and preference hierarchyPowerful but hard to inspect and update.
Deployment-time YAMLVersioned, diffable rules for pipeline mode, risk mode, allow/deny, budgets, and audit destinationsDirectly inspectable by non-model engineers.
Programmable policy languagePredicates, quantifiers, automata, or UI transition DSLsMore formal power, less portability.

The open question is composition: can deployment-time governance reliably override training-time dispositions, and can policy schemas become portable across harnesses?

Audit Infrastructure

Governance requires accountability. A replayable audit record should include trace IDs, principal identity, tool calls, policy decisions and versions, execution results, resource costs, and integrity hashes for relevant inputs/outputs.

Most systems record only a subset. Few sign or hash enough of the record to protect it from a compromised process.

The paper distinguishes:

  • per-action monitoring: cheap and local, but misses slow multi-step attacks;
  • trajectory-level audit: better for multi-step patterns, but higher latency and harder triggering.

Practical systems often run per-action checks inline and trajectory-level analysis asynchronously over audit logs.

Governance Coverage Is Sparse

The survey’s governance coverage matrix encodes representative systems across permissions, hooks, hardening, constitutions, audit, and multi-agent governance. The important result is that no system covers the whole surface.

System familyStrengthGap
Codex / Gemini CLI / OpenHandsSandboxes, action approval, partial auditFine-grained identity and formal policy portability.
AutoHarnessDeclarative risk-tiered governanceSchema portability and external interoperability.
Progent / ConsecaContext-dependent permissions and generated policy checksGeneral runtime enforcement and standardization.
CaMeL / SAFEFLOWInformation flow and transaction semanticsIntegration cost and flexibility tradeoffs.
SAGA / Authenticated DelegationIdentity, delegation, short-lived credentialsEnd-to-end integration with tools, traces, and audit.
AgentSpec / AgentDoGProgrammable hooks and trajectory diagnosisHook standards and compositional effects.

Governance is therefore not a security toggle. It is a cross-layer control plane spanning identity, tools, context, lifecycle, audit, and human authority.

Security Landscape

The paper maps governance mechanisms to risk categories from the broader agent-security literature:

  • untrusted interfaces;
  • wrong instruction following;
  • unconstrained data flow;
  • hallucination;
  • data leakage;
  • unauthorized action;
  • resource exhaustion.

It also notes contextual security as a proposed extension beyond the CIA triad. That framing is useful, but the paper treats it as an emerging proposal rather than an established standard.

Research Directions

Governance-specific open directions include:

  • standardized policy and audit languages;
  • formal guarantees over governance pipelines;
  • adaptive governance and governance of policy generators;
  • long-horizon governance with renewal, revocation, and audit over evolving sessions;
  • usable permission UIs, audit dashboards, and constitution editors;
  • cross-layer coherence between training-time, deployment-time, and runtime governance;
  • end-to-end supply chain governance;
  • unified adversarial benchmarks for full governance stacks.
Was this page helpful?