Selected work / Continuity & evaluation

Keep the understanding, not just the history.

Long-running collaboration accumulates decisions, corrections and reasons. I built cdx-continue to preserve what makes the next step intelligible, and code-bench to compare agent setups through the work they actually produce.

Role
Creator · product direction and evaluation design
Scope
Continuity · evaluation · collaboration
State
Working personal tools and design explorations

Remembering facts is only part of continuing

An AI conversation eventually needs a smaller working context, or the work moves to a fresh conversation. A summary can keep the files and completed tasks while losing why a decision was made. The successor then repeats a rejected approach, asks a settled question or quietly narrows the original ambition.

The lost material is often a relationship: a requirement with a condition, an experiment deferred rather than abandoned, a correction that changes how later results should be read. Creative work and research depend on that meaning just as much as software does.

I separated purpose from current state

In cdx-continue, the enduring brief carries purpose and the conditions of success. A separate continuation document carries current understanding, unresolved questions and the next useful action. I chose that separation so a temporary obstacle does not become a permanent instruction.

The document also preserves consequential effects. “The message was sent” and “the message is ready to send” cannot collapse into the same note. Continuity changes what an agent does in the world, not only what it recalls.

I pushed the design away from an ever-growing log. Settled history becomes its surviving meaning; details remain reachable where they belong. The next agent should recognise what matters without reconstructing the whole conversation.

The receiver is the real test

A summary can look excellent to the agent that wrote it, which still knows the original conversation. I test what a fresh receiver can understand without coaching: the purpose, an accepted choice, a consequential limit, a commitment waiting for the right condition.

A new version need not win every comparison, but it should preserve what matters at least as well as the existing one. The comparisons behind the current version supported that choice while also revealing an omission both versions shared and increased context use. That is evidence for a particular choice, not a promise of perfect memory.

From a promising prompt to a controlled comparison

Code-bench asks whether my actual agent setup improves coding results, and at what cost. Paired runs compare a pristine setup with my local setup, or native continuation with a fixed cdx-continue configuration. Matching tasks, models and relevant settings makes those differences interpretable.

The result keeps correctness beside resource use and activity. I want to see the trade-off, not reward an agent for looking busy.

More ways to collaborate

Continuity also matters when work crosses from one participant to another. These projects explore three different relationships: correspondence, a delegated contribution and a meeting of perspectives.

Message: correspondence with consequences

Message is a private local correspondence hub for a person and explicitly authorised agents. Find the recipient, read the conversation, send or reply, then inspect the outcome. I chose one shared history and outbox for the Mac app and command line, with explicit destinations rather than hidden routing. It is installed and in use in my own environment, and still under development.

Grokpair: an independent contribution

Grokpair lets an agent hand a bounded research, review, implementation or creative task to a detached Grok Build collaborator while continuing its own work. The caller reads the returned answer and artifacts and decides what to adopt. I kept integration and responsibility with that caller: another perspective is useful input, not automatic authority.

The Seven: exploring collective judgment

The Seven is a design exploration: six models answer independently, each evaluates the candidate answers, and a seventh synthesises a response. The idea is an answer-first experience with the underlying contributions available to inspect. Its current platform direction is documented, not implemented or deployed; whether that arrangement yields a better answer remains the question to test.