AI is easy to call. Making it fit into real work is harder.

New models and services appear without pause, and tasks that were difficult yesterday can sometimes be completed today with a few lines of conversation. But more capability also means more choices: which model should see the work, which context should it receive, what may it change, and how do I know the result is safe to continue with?

I wrote this post as the starting point for exploring those questions. I am less interested in presenting AI as a magic button than in building small tools that help people carry their own work forward. The useful result is not the most impressive response. It is the next step becoming easier to review and take.

The real problem is fit

A capable model can still be inconvenient. People have to repeat the same background, sort a large answer before using it, or worry that a helpful connection also gave an agent access to too much. Convenience is therefore not just a user-interface problem. It includes context, authority, feedback, and recovery.

A human developer moves from context to an AI suggestion, reviews the result, and continues the work in one connected loop

A useful AI tool keeps the person in the work loop instead of hiding the important decisions.

The tools I want to build should do four modest things well:

Question Design direction
What should the agent see? Give it a small, relevant context projection instead of an unrestricted history.
What may it change? Make the authority and stopping point explicit before a tool call.
How do we know it worked? Inspect the artifact, evidence, and validator result rather than trusting a fluent answer.
What happens when it fails? Return a bounded failure and keep the human workflow recoverable.

This is why the later work around Luthn and Drig became more concrete than the original idea of “making AI easier.” Luthn separates safe context from protected source material. Drig separates planning and post-development QA from ordinary product implementation. Those are not features I can use to claim that every workflow is now better. They are boundaries that make a workflow easier to understand and improve.

What I have learned while building

The first lesson is that more automation is not automatically more convenience. A long chain of planning, delegation, retries, and summaries can create more state than a person can inspect. A small task often benefits from a narrow tool surface and a short feedback loop.

The second lesson is that memory needs permission. Remembering a preference or a project decision is useful only when the system can say where it came from, who may see it, and when it stops being valid. A connection to a memory service is not a permission to read every record behind it.

The third lesson is that the harness is part of the product. The model, prompt, tools, context, retry policy, and validator together produce the result. If any of those conditions are hidden, a success is difficult to reproduce and a failure is difficult to diagnose.

These are observations from the public development work documented here, not a controlled productivity study. I am not claiming that one architecture or one model wins for every team. The claim is narrower: tools become more useful when they make context, authority, evidence, and recovery visible.

What this blog will record

Future notes should show more than a finished screenshot. When the evidence is available, I will record the environment and time, the action I took, the result, the failure or surprise, and the limit of the conclusion. When the evidence is not available, I will label a statement as a source-reported claim or leave it as an open question.

That is the standard I want this blog to follow. AI should feel easier because the surrounding work is clearer, not because uncertainty has been hidden. The tools may start small, but each one should leave a person with a result they can inspect and a next step they can choose.