Archive
All posts
41 notes, newest first.
All notes
- I Went Back to Codex, Then Built Another HarnessI split coding and assistant work between Codex and Hermes, then built Drig to orchestrate the whole workflow. The coordination and QA overhead eventually brought me back to one tool.

- I Lost to the Harness First: My OpenCode AI FailureA developer retrospective on leaving GitHub Copilot for OpenCode, then quitting a multi-agent harness when token cost outran the product work.

- Hermes, Grok Bot, and Buzz: What a Single AI Agent Comparison MissesA comparison of Hermes, Grok Bot, and Buzz based on early hands-on use and current primary sources. The useful distinction is not model quality, but automation scope, access, and approval boundaries.

- Is Software Ending? Notes From Building With AI AgentsA project retrospective on writing less code by hand, trusting AI agents more, and why responsibility may become the developer's real job.

- When the Sandbox Breaks, What Does Luthn Actually Stop?An inspection of Luthn’s memory, sensitive-data, and authority boundaries after a frontier-agent intrusion made sandbox assumptions harder to trust.

- Security Is About Designing Authority, Not Just Hiding DataA development retrospective on turning sensitive AI access into bounded requests, safe projections, explicit approval, expiry, and testable failure behavior.

- Dograh: Open-Source Voice Agents You Can Run YourselfA source-backed look at Dograh's self-hosted voice-agent stack, the control it promises, and the operational questions an adopter still needs to answer.

- Can Korea Turn Semiconductor Strength into AI Products?A look at Motif 3, K-EXAONE 2.0, A.X K2, and Solar Open 2—and what Korea needs to turn model scores into AI products.

- Cursor Is Now Part of SpaceX. Will Developers Keep Their Choice?Cursor now has SpaceX compute behind it. The next test is whether a stronger workbench can preserve meaningful model, cost, data, and operational choice.

- OpenAI Wants Safety Signals Without Keeping Your DataPrivate Safety Processing is an early OpenAI preview for finding misuse patterns across related interactions while remaining compatible with eligible Zero Data Retention deployments.

- Can Higgsfield Become the Studio Behind Your Next AI Video?An introduction to Higgsfield's camera-controlled AI video workflow, recent product and funding news, pricing, market potential, and limits.

- More Agent Memory Does Not Always Make an Agent BetterResearch and Luthn’s bounded recall design show how agent memory can improve reliability while narrowing exploration and increasing context cost.

- Stripe’s OpenRouter Deal Puts Model Routing at the Center of AI InfrastructureStripe's reported OpenRouter acquisition shows why model selection, usage tracking, and AI billing are becoming one infrastructure layer.

- Grok 4.6 Is Moving Fast. How Far Will Cursor Change?What SpaceX’s Cursor acquisition and the official Grok 4.6 launch could mean for Cursor as an AI agent workbench.

- Why I Stopped Automating Development and Started Enjoying Coding AgainA project retrospective on narrowing Drig from an all-in-one development loop to bounded planning, evidence, QA, and delivery around a direct Codex workflow.

- How Far Have Open-Weight Models Entered Real Work?Open-weight models are becoming a practical choice for building local agents and task-specific AI stacks. Meta Muse Glimmer and DeepSeek V4 Flash show both the opportunity and the limits.

- Lovable: The AI Cofounder That Turns Ideas into Web AppsA practical introduction to Lovable for non-technical founders and teams, covering its workflow, pricing, market potential, and limits.

- Claude's Watermark Reveals Two Directions for AI EcosystemsAnthropic's invisible watermarking and account boundaries raise a broader question about controlled AI services versus inspectable developer tools.

- Agent Memory Is Harder to Audit Than to StoreLong-term memory research points beyond storage and retrieval to validity, relations, provenance, and auditability, with Luthn as a practical safe-context example.

- The Damage Vibe Coding Is Already Causing in KoreaVibe coding lowers the barrier to building software, but the costs of operations, security, and accountability do not disappear.

- Needle 2 and the Possibility of LLMs Built Into More ThingsNeedle 2 shows how a small on-device model can turn the destination, tool contract, confidence threshold, and escalation path into part of an LLM's design.

- When Tool Calls Become Patent TerrainMistral’s U.S. patent on code-implemented tool calls has triggered a broader question for agent developers: which parts of a tool-calling runtime are generic practice, and which implementation details deserve a closer IP review?

- Meta Chooses 30B for Local Agents with Muse GlimmerMeta's Muse Glimmer makes local agent execution a concrete hardware choice, while its memory, runtime, and privacy boundaries still belong to the surrounding system.

- When Safety Testing Changes the AI Release CalendarRecent cyber evaluations show that isolation, monitoring, and release gates are becoming part of frontier-model development, not just post-launch security.

- An AI Agent Needs an Identity, Not Just an AccountAs agents act on behalf of people, identity, delegated authority, and verifiable audit trails matter more than simply issuing another credential.

- Agent Memory Needs a Permission Model, Not Just SearchWhat building Luthn taught me about owner isolation, safe projections, protected-access workflows, and the tests that must run before retrieval tuning.

- The Browser Is No Longer Just for HumansA comparison of Playwright, Computer Use, Browser Use, and Cloudflare Browser Run—and why the browser is becoming an execution environment for AI agents.

- The Missing Variable in Agent Evaluation: The HarnessHarness-Bench shows why agent results should be reported with the execution configuration that supplies context, tools, permissions, recovery, and verification.

- Local LLM Agents Need Memory Audits, Not Just More ContextAs smaller open-weight models move onto personal hardware, long-term memory classification, provenance, and auditing become core agent design problems.

- As AI Slop Grows, Better Sorting Matters More Than More SearchAI-generated volume is changing the signal-to-noise ratio of technical information, making provenance and verification central to how developers read and remember.

- Codex Updates Are Changing the Work Surface, Not Just the ModelOpenAI's recent Codex updates connect GPT-5.6 access with browser debugging, remote work, and reusable workflows.

- What Humans Miss When Approving AI Agent CommandsA 40,000-run experiment shows why familiar commands, approval fatigue, and human-in-the-loop controls are not enough for coding agents.

- Why I Do Not Give One Codex Model Every JobA project retrospective on separating planning, implementation, review, and QA by responsibility, model route, evidence, and permission rather than simply adding agents.

- Can Google's Ecosystem Become an AI Growth Engine?Google has an extraordinary AI ecosystem. The harder question is whether that ecosystem can create the focus and market momentum that OpenAI and Anthropic have built around their AI products.

- GLM-5.2 and the Safety Gap in Open-Weight AIA look at GLM-5.2's performance gains and the safeguards, evaluation, and deployment challenges that independent testing has highlighted for open-weight models.

- AI Agent Actions Should Be Recorded in External Audit LogsA look at the basic principles of external audit logs, permission boundaries, and real-time detection as AI agents gain access to real systems.

- The Problem with AI Agents Is Permission Design, Not IntelligenceThe recent cybersecurity evaluation incidents show why agent safety depends on authority boundaries, not only on prompts, model quality, or a larger harness.

- How AI Regulation May Spread After the EU AI ActAn outlook on the EU AI Act's implementation, likely responses in the United States and South Korea, and a checklist for developers preparing to deploy AI services in the EU.

- Open Source Does Not Grant Permission to Use Someone Else's DataThe K-skill and Blue Ribbon controversy illustrates the difference between an open-source license and permission to use external data, as well as the responsibilities of developers in the age of vibe coding.

- GPT-5.6 Luna Gets an 80% Price Cut: OpenAI's Next MoveOpenAI has lowered API pricing for GPT-5.6 Luna and Terra. Here is what changed, how developers are responding, and where this shift may be heading.

- Building Tools to Make AI Easier to UseA starting note about building AI tools around real workflows, bounded context, explicit authority, and results that people can review.

No posts match those filters.