Personal assistant ops architecture

YAAA

an assistant you actually own

Yaaa (Yet Another AI Assistant) is a personal assistant ops architecture: a self-hostable, local-first system where memory, routing, and every action stay versioned, reviewed, and owned by you. Sensitive work runs local; cloud and model use is bounded by mode routing and leak checks that keep sensitive context off disallowed routes, and side effects leave through one fail-closed gate. Conversation harnesses can be replaced without moving the source of truth.

Amanuensis is one concrete implementation within this framework, supplying reasoning, converse, and attention-aware delivery subsystems for information triage. Yaaa defines the ops architecture around those subsystems rather than replacing them with one application.

The more context a personal assistant needs, the more consequential its ownership boundary becomes.

Context crosses a boundary

To be useful, an assistant needs context other software only glimpses: identity, conversations, private notes, location, and sometimes legal or medical history. Sending that context across an external ownership boundary makes provider controls consequential, and a complete recall path may not be available.

Provider governed

Data routed through a cloud service may be subject to provider logging, retention, moderation, and policy. Those terms may be reasonable for some tasks, but they are not controlled by the user.

Locked · fragmented

Each vendor keeps agent memory in its own format, and nothing syncs between them. Memory drifts across ChatGPT, Claude, and the next tool; you reconcile the gaps by hand, and human memory fails too. An assistant meant to lift that burden quietly adds to it.

Yaaa starts from the opposite end: an assistant for you, on infrastructure you own, where memory is one reviewable record, sensitive work stays local, and cloud and model use is bounded by mode routing and leak checks. Those checks stop sensitive context before it reaches a disallowed model or channel. Ownership is not a feature toggle. It lives in the architecture.

What it is

A self-hostable assistant architecture

Yaaa is an opinionated multi-agent architecture for the tools you already run: mail, notes, calendar, home systems, local models, and optional cloud models. It turns those separate runtimes into one assistant by giving them shared contracts for memory, routing, action, and governance.

Boundary

Replace parts without moving authority

Conversation harnesses, models, tools, and source adapters can change. The authority path stays fixed: reviewed memory as the source of truth, local-first routing for sensitive context, one write boundary for side effects, and a governance loop for changes that affect future behavior.

The design question is ownership: where memory lives, which runtime may see sensitive context, and what must happen before an action leaves the system.

You keep the steering wheel. Yaaa carries the routine work: gathering context, routing the task, keeping the record, and asking before anything reaches the outside world.

Ops discipline, different subject

LLMOps

Runs models in production for a team: serving, prompt and fine-tune pipelines, monitoring, evaluation, throughput, and cost. It optimizes how well an LLM service operates.

Yaaa

Runs one assistant for one person: a reviewable memory of record, sensitivity-aware routing, gated actions, and portability across harnesses. It optimizes who owns and can audit the result.

Same rigor: versioning, review, gates, audit. The subject is ownership rather than throughput.

The high-level design is broad read, local-first binding, bounded reasoning, and one write boundary. L0/L1 gather context; L2 chooses mode, model, and tools; L3 works the task; L4 decides whether any outside-world effect may run.

Yaaa separates two loops. The runtime loop handles the current task. The governance loop studies traces later, promotes reviewed memory, SOPs, ADRs, runbooks, and policy changes, then applies them back to future runs. That keeps the assistant adaptive without letting a model silently become the authority.

Layer breakdown

  • L0IO boundaryNames how outside events enter and how approved outputs leave: passive input, proactive input, passive output, and proactive output.
  • L1SENSETurns available context from adapters into source records with a unified shape and preserved provenance.
  • L2Binding layerBinds memory authority, local-first mode routing, model/tool choice, and the replaceable Converse surface.
  • L3Bounded ReActRuns the task loop under L2 context and returns observations or an intent proposal, never direct writes.
  • L4Gated actionConverts side-effect intent into approve, refuse, or defer through one policy-checked write boundary.
  • L5Governance foundationTurns run traces into reviewed memory, SOP, ADR, runbook, manifest, or policy changes for future runs.

Architecture diagram

Runtime flow moves through L0-L4 for the current task. L5 sits below it as the reviewed feedback path that can change future behavior.

  • ENTITYrecords and artifacts
  • FLOWtransforms
  • ACTIONside effects
  • CTRLpolicy and gates
L0 - IO boundaryL1 - SENSEL2 - Binding layerL3 - Bounded ReActL4 - Gated actionL5 - Governance foundationpassive inputENTITYpassive inputRSS, feeds, timersproactive inputENTITYproactive inputoperator turn / approvalpassive outputENTITYpassive outputinbox / trace viewproactive outputACTIONproactive outputpush / ask / alertSENSEFLOWSENSEsource records insource adaptersFLOWsource adaptersfeeds, docs, eventsnormalizeFLOWnormalizeparse, classifysource recordsENTITYsource recordsoperator-owned copymemory SSoTENTITYmemory SSoTgit-backed authoritypolicy routerFLOWpolicy routermode/model/tool routingharness memoryENTITYharness memoryrebuildable projectionagent memoryENTITYagent memorystaged namespaceconverseFLOWconversereplaceable session harnessReAct task loopFLOWReAct task loopcalled with contextreasonFLOWreasonplan next stepactACTIONactintent proposal onlyobserveFLOWobserveresult / tracereflectFLOWreflecttask-local adjustfail-closed gateCTRLfail-closed gatedefault denyside-effect intentENTITYside-effect intentproposal, not authorityprocedure manifestENTITYprocedure manifestversioned write planpolicy checkCTRLpolicy checkmode + permissionapprove / refuse / deferACTIONapprove / refuse / deferrun, stop, or holdgovernance meta-loopCTRLgovernance meta-loopauthority change reviewrun tracesENTITYrun traceswhat happenedextractFLOWextractcandidatesdistillFLOWdistillstructured artifacthuman reviewCTRLhuman reviewapprove / rejectpromoteACTIONpromoteapproved changememory reconciliationFLOWmemory reconciliationreviewed namespace alignmentapplyACTIONapplyreviewed runtime behaviorADR changelogENTITYADR changelogdecision historyrunbook updatesENTITYrunbook updatesSOP + rollback
Yaaa L0-L5 system map: IO boundary, SENSE, binding layer, bounded ReAct, gated action, and governance foundation
Layer detail
View

Double-click a layer to show detail. Use Fit to focus the active detail.

High-level view

Architecture drill-in

Deep dive into the layers

The full map shows the contract; the drill-in names what each layer owns, what it may see, and where its authority stops. Read it as an opinionated guideline: L0/L1 gather context, L2 binds memory and routing, L3 proposes, L4 controls side effects, and L5 changes the rules future runs stand on.

L0IO boundary

L0 is the system's contact boundary with the outside world. Inputs enter and outputs leave through named paths, so model access and channel delivery stay on the record.

The four quadrants name who starts the event and whether it enters or leaves:

  • Passive inputRSS, docs, timers, and subscribed sources arrive for SENSE.
  • Proactive inputYour prompt, reply, correction, or approval enters through converse.
  • Passive outputAn inbox, trace view, or project surface waits for you to check it.
  • Proactive outputAn ask, alert, or result reaches out through the routed output path.
passive inputproactive inputpassive outputproactive output

initiator + direction define the route

L1SENSE

L1 turns ambient input into durable source records. Source adapters connect to each source; normalize parses and classifies into one shape; the result keeps provenance attached.

Example: an RSS article arrives overnight. SENSE stores the feed URL, fetch time, source title, normalized excerpt, and classification, then keeps that record ready for later reasoning.

An owned record means the deployment keeps a local, auditable copy or pointer under your control, with source identity preserved. Durable memory is a later L2 promotion after review.

source adapters
connect and pull each source
normalize
parse and classify into one shape
source records
user-owned, provenance preserved
source adaptersnormalizesource records

available context aggregated into a unified format

L2Binding layer

L2 keeps memory as three surfaces with different authority. The memory SSoT is the git-backed record that accepts reviewed changes. Harness memory can grow inside each conversation surface as local recall or serving state. Agent memory holds staged, task-local deltas.

Routing lives here too: the policy router reads the task mode, sensitivity, and locality policy, then selects the model, tool, and read accessor. A health question or private calendar change stays on the local route; a low-sensitivity lookup may use a scoped cloud model if the mode allows it.

The router points the selected path toward Converse, the replaceable session surface. For example, the same request can arrive from a web chat, terminal prompt, or voice shell; Converse packages the turn, receives the reply, and hands the routed work forward while memory authority and write authority stay in L2/L4.

memory SSoT
git-backed authority: the record
harness memory
local recall / serving state
agent memory
staged, task-local deltas
policy router
mode / model / tool routing
converse
chat / terminal / voice turn surface
memory SSoTharness memoryagent memoryconversepolicyrouter

Converse packages a web, terminal, or voice turn

L3Bounded ReAct

L3 is a called task loop. L2 supplies memory and routing context; the loop returns observations or a side-effect intent for L4.

The load-bearing constraint: "act" emits an intent proposal only. The loop can plan, try, observe, and adjust while every outside-world effect waits for L4's gate.

reason
plan the next step
act
proposes an intent only
observe
read the result / trace
reflect
task-local adjustment
ReAct task loopreasonactreflectobserve

L3 loops on the task; act emits intent only

L4Gated action

L4 is the system's one write boundary: the place where a proposed side effect can become a real one. The design keeps all outside-world mutation in one auditable path, so every harness follows the same checks for target, account, permission, confirmation, sensitivity, and failure behavior.

An action arrives as a side-effect intent, which is a request. It expands into a versioned procedure manifest; the policy check weighs mode, permission, and procedure constraints; then the fail-closed gate defaults to no execution. A missing manifest, ambiguity, an expired confirmation, a timeout, or an unknown action class resolves to "do not run". The outcome is one disposition: approve, refuse, or defer (held safely, unexecuted; resuming re-enters the gate).

One boundary also makes replacement cheap: a new chat surface, local runner, or model adapter can propose intent, but it inherits the existing write contract instead of carrying custom authority.

side-effect intent
a proposal awaiting policy
procedure manifest
the versioned write plan
policy check
mode + permission constraints
fail-closed gate
default deny
approve / refuse / defer
run, stop, or safely hold
side-effect intentprocedure manifestpolicy checkfail-closed gateapprove / refuse / defer

the write funnel

L5Governance foundation

L5 is called the foundation because it sits under runtime authority. It is the reviewed path that changes the rules future runs stand on: memory, policy, SOPs, ADRs, runbooks, manifests, and delivery behavior.

Run traces from conversations, tool calls, procedures, and failures are extracted into candidates and distilled into structured artifacts. A human review gates any authority change. Approved deltas are promoted into the relevant source of truth; memory reconciliation aligns the SSoT, harness memories, and agent namespaces; apply means a reviewed artifact can change future behavior.

Example: after repeated appointment reschedules, traces show the same timezone and confirmation checks. L5 can promote a runbook update and a stricter manifest requirement, so the next run starts from reviewed procedure rather than rediscovering the rule.

run traces
raw evidence from runs
extract / distill
candidates into structured artifacts
human review
the gate before any authority change
promote
land approved deltas into the record
memory reconciliation
align SSoT, harnesses, namespaces
apply
a reviewed change alters future behavior
run tracesextractdistillreviewpromotereconcileapply

L5 turns traces into reviewed future behavior

Passive sources may arrive on timers or subscriptions. Interactive work enters through Converse. Reads skip the write gate but still carry egress risk, so high-sensitivity reading requires mode-aware accessors and leak checks that stop sensitive context before a disallowed cloud route.

There is a second loop above the first: observe runs, extract candidates, distill them into artifacts, review by a human, promote into the record, reconcile across harnesses, deploy. The assistant learns by turning runs into auditable artifacts rather than silently mutating opaque model memory.

Invariant 01

Memory that stays yours

Memories can grow where work happens. Each harness may keep local recall, task notes, and staged deltas. Yaaa keeps a git-backed SSoT as the conflict ledger. Harnesses contribute changes up and pull reviewed changes down; reconciliation records conflicts and resolutions.

Federated memory can grow; authority stays reviewable.
Invariant 02

Traceable decisions

Repetitive work starts from reviewed decisions, SOPs, runbooks, and procedure manifests. They give the agent the approved shape of the work: what is known, what is allowed, what must be checked, and where the risk boundary sits.

The model reasons against recorded decisions before it proposes intent.

Less guessing, more replayable intent.
Invariant 03

Fail-closed actions

Outside-world effects pass through one deterministic gate. If the assistant wants to reschedule an appointment, it can draft the calendar change and message, but L4 must see a manifest: recipient, account, timing, text, sensitivity, and required confirmation.

If the recipient is ambiguous or approval has expired, the gate defers or refuses; nothing is sent quietly.

The assistant proposes; the gate decides.

The personal assistant triangle

Yaaa makes the tradeoff explicit

Personal assistants pull against three desires: deep capability across closed ecosystems, low-friction automation, and owned, auditable control. Yaaa chooses the control corner first. The consequence is narrower integration, more review at sensitive boundaries, and a system that is honest about what it cannot see or safely do.

How much friction does safety add?

Yaaa gates outside-world side effects, not every answer. Reading, drafting, and local reasoning can stay fast; L4 is reserved for actions that send, change, spend, publish, or otherwise affect the world. Stable low-risk procedures can graduate into narrow policy grants, where the gate records an approve_by_policy disposition instead of asking every time.

High-risk or ambiguous actions still slow down. A sensitive action that cannot name its target, account, permission, confirmation, and rollback path should defer or refuse instead of running by accident.

Where is Yaaa intentionally slow?

The slow path is the sensitive path: provenance-preserving ingestion, local model routing, side-effect review, and later governance. L0-L5 are authority boundaries, not one mandatory synchronous chain; passive input can be sensed ahead of time, L5 runs after the task, and low-risk reads can use a shorter route.

Yaaa is a poor fit for millisecond-critical experiences such as live translation or always-on voice control. Local hardware, adapter quality, review steps, and retention budgets set real speed limits.

Can natural-language memory really be reconciled?

Runtime memory should live where runtime systems are good at it: Postgres for source records, traces, working memory, and queues; vector indexes for retrieval; harness and agent memory for local recall. The git-backed SSoT holds reviewed authority: core memory, SOPs, ADRs, policies, and procedure manifests.

LLM-assisted reconciliation can explain diffs and propose candidate patches, but the model is not the judge. Semantic conflicts still need policy or human review before one version becomes authority.

What happens when a data source is closed?

SENSE treats adapters as graded capabilities: official APIs and exports first, local files, webhooks, RSS, and manual capture next, with RPA or OCR only as a fragile fallback. Missing sources stay explicit, so the assistant can say what it knows and what it cannot see.

Closed ecosystems remain closed. RPA can help the system see, but it is slow, brittle, less auditable, and vulnerable to UI or platform changes, so it should not become a high-trust source.

Who is willing to operate this?

The operator should not live in YAML. Yaaa needs opinionated default policy sets, plain-language review queues, and one-click promotions for stable lessons. L5 should turn traces into simple decisions, such as whether a repeated calendar pattern should become a default rule.

This remains a self-hostable architecture for operators and power users first. A zero-admin consumer assistant is a different product boundary, and some maintenance burden is part of the bargain.