Problem & my contribution
A regular AI chat has limited memory and little connection to real work. Giving it tools creates another problem: crossing project boundaries, repeating an action after a failure or treating its own inference as permission.
My contribution. I define the assistant’s roles, interaction rules, memory boundaries and action policies. My responsibility is the product logic and system design: what it may decide, when it must ask and how the user understands what happened.
A coordinator delegates to specialist roles. History stays scoped to its owner and conversation. Sharing a fact with a project requires an explicit decision. Risky actions pass through a separate approval queue.
Outcome
The source implements shared browser/iPhone conversations, voice input and spoken replies, message history with actual role attribution and in-chat approvals. These are delivered capabilities; no unmeasured time savings or user counts are claimed.
Engineering highlights
Do not repeat an unknown outcome
Losing a response does not prove an action failed. Provider fallback is restricted to eligible failures before execution.
Memory with provenance
Facts reference source messages. Ownership, revision checks and writes share a transaction, so a stale model response cannot overwrite newer memory.
Delegate only when useful
The coordinator selects relevant roles. Call limits and cycle detection keep delegation bounded.
Architecture
| Layer | Technology & purpose |
|---|---|
| Interfaces | Telegram, Preact, SwiftUI; browser chat |
| Backend / AI | Bun, TypeScript, role orchestration, Claude Agent SDK |
| Data | SQLite WAL; conversations, tasks, approvals, project memory |
| Infrastructure | systemd, GitHub Actions, separate execution bridge |
Trade-offs & lessons
SQLite keeps operations simple and memory updates atomic, but long operations in one process need care. Voice uses recording → transcription → response → synthesis: easier to control, but not full-duplex speech.
What the implementation taught
A documented dashboard defect counted “not paused” as “online”, allowing 12/12 above red indicators. The fix made the counter and indicators use the same health predicate. Lesson: process state and readiness to do work are different facts.
What I would improve now
I would start with one executable scenario set across Telegram, browser and iPhone, then measure response latency and time to an approved action. This is a next step, not a claimed result.
Operations & data
The service and dashboard are deployed separately from the signed iPhone build. Queues, action logs and health states distinguish completed work, pending decisions and failures. Credentials and production databases are excluded from the public repository.
Data & migrations
A conversation belongs to an owner; project membership is a separate mapping. Chat memory carries a revision, proposals snapshot it, and approved knowledge preserves provenance. New tables are additive to the existing conversation store.
| Layer | Relationship / rule |
|---|---|
| Conversations | owner → conversation → messages |
| Memory | conversation → revision + sourceMessageIds |
| Projects | project → approved knowledge |
| Approvals | proposal → owner decision → action |
Limits
Provider accounts and owner setup are required. Working conversations are not public demos. The 2D/3D office and GitHub MCP adapter remain on a development branch; banking execution is not presented as a shipped capability.
Evidence & source
Verified 24 Sep 2026: 15 memory/fallback checks; dashboard build. These are targeted checks of the described decisions, not a full application audit.
Claims above are tied to public code and documentation. Evidence links point to the reviewed revision; CI status may change.

