The garden
Unlike a blog, a digital garden isn't organized by date and isn't published "finished." Notes are planted rough, tended over time, and linked as connections emerge. The maturity marker tells you how much weight to put on each one. 267 notes and growing — fed daily by a research knowledge base of 32,000+ concept notes.
Start here
Choose a question, not an archive.
Explore an architecture territory.
Currently tending
A few notes to begin with.
Agentic AI
134 notes
-
Designing memory for agent systems
Agent memory is a data modeling problem wearing an AI costume — and the 2026 memory-system landscape (Letta, Cognee, bi-temporal organizational graphs) has converged on the same answer: typed, structured, provenance-carrying records. Schema decisions matter more than embeddings.
Aug 2026 -
Agent audit trails must answer: who acted, for whom, with what authority
"The agent did it" is not an audit answer; regulated environments need an accountability chain designed into the architecture, because it cannot be reconstructed afterwards.
Sep 2026 -
Agent memory needs governance before it needs embeddings
Persistent memory is an attack surface and a compliance surface; retention, provenance, access control and deletion must be designed before retrieval quality matters — and memory control-flow hijacking makes the store a command channel if you don't.
Aug 2026 -
Agent observability means trajectory replay, not dashboards
Metrics dashboards answer how often an agent fails; only full trajectory capture and replay answers why — the unit of observability for an agent is the complete decision trace, and replayability should be a design requirement.
Aug 2026 -
Agent reliability is failure-mode design, not uptime
An agent can be 100% up and still failing constantly; reliability engineering for agents means enumerating and designing for semantic failure modes, not infrastructure availability.
Aug 2026 -
Agents need checkpoints the way databases need write-ahead logs
Long-running agent work must survive process death, redeploys, and rate-limit stalls; durable execution is the WAL of agentic systems — an architectural requirement, not an optimization.
Aug 2026 -
Choose workflows where determinism is cheap, agents where it's expensive
The workflow-vs-agent decision is economics: if the path can be enumerated in advance, a deterministic workflow is cheaper, testable, and auditable — spend agent autonomy only where enumeration costs more than supervision.
Aug 2026 -
Context compaction is lossy compression — choose your loss function
Every compaction strategy is a lossy compressor; the failure isn't compression, it's never deciding what class of information you're willing to lose.
Aug 2026 -
Context engineering is the new schema design
What goes in the context window, in what structure, with what priority, is schema design — representation, grain, and ownership of information. Prompt engineering by vibes is schema design done badly.
Aug 2026 -
Cost is an architectural property
LLM cost is determined by architecture decisions — caching strategy, context discipline, model routing, gateway budgets — made months before the invoice arrives.
Sep 2026 -
Evals are the contract, not the report card
Evals are the executable definition of "working" that gates every change, like a test suite; without them every prompt tweak and model upgrade is an unreviewed production change.
Aug 2026 -
Hallucination is a systems symptom, not a model defect
Production hallucination usually traces to the pipeline — failed retrieval, stale knowledge, ungrounded prompts, missing abstention paths — so treat it with systems engineering, not model shopping.
Sep 2026 -
Knowledge staleness is a data-lifecycle problem
Freshness is a lifecycle contract over observed, effective, ingested, and verified time—with query-class SLOs and explicit fallbacks when evidence is too old.
Sep 2026 -
Long-horizon agents don't crash — they lose the plot
The characteristic failure of long-running agents is gradual coherence loss that no retry logic catches, because every individual step succeeds; you need budgets, drift detection, and deliberate re-grounding checkpoints.
Aug 2026 -
Memory writes are schema decisions; memory reads are query plans
Retrieval quality is mostly decided at write time. Treat agent memory reads the way a database treats queries: plan them, don't just embed and hope.
Aug 2026 -
Multi-agent systems fail along their communication edges
Failure taxonomies keep finding that individual agents mostly do their jobs — systems fail at the edges: misunderstood handoffs, divergent state, lost context. Topology and edge contracts deserve the design attention.
Sep 2026 -
Prompt injection is a permissions problem, not a prompting problem
You cannot instruct your way out of an architecture flaw. Assume injection succeeds, then bound what it can reach.
Aug 2026 -
RAG is a data pipeline with a retrieval step
Most RAG failure is ingestion failure — chunking, freshness, metadata, index lifecycle — while teams debug the retriever and reranker downstream.
Aug 2026 -
Session state and long-term memory are different schemas
Conflating resumable session state with long-term memory produces systems that do both badly — the lifecycles, consistency needs, and storage are different.
Aug 2026 -
Size the sandbox to the blast radius
Sandbox selection is a blast-radius calculation, not a security fashion choice: ask what the worst accepted action costs, then buy exactly that much isolation.
Aug 2026 -
Slowly changing dimensions already solved memory versioning
The warehouse world solved 'a fact changed and we need both current truth and history' decades ago; agent memory teams are reinventing it badly.
Aug 2026 -
Structured outputs are the integration contract
Constrained decoding establishes shape, semantic validation protects meaning, and measurable repair remains the final fallback at the model-to-software boundary.
Sep 2026 -
Sub-agents are context isolation, not delegation theater
The engineering value of sub-agents is a fresh, scoped context window — isolating noisy exploration and bounding blast radius — not an anthropomorphic org-chart of specialists.
Aug 2026 -
The demo-to-production gap is an architecture problem
80% of enterprise AI initiatives stall between demo and production. The model was never the problem — integration, data, and ownership are.
Aug 2026 -
The harness is the product; the model is a component
The differentiated engineering in an agent system lives in the harness — context assembly, tool execution, permissioning, loop control, recovery — while models are swappable components.
Sep 2026 -
Tool design is API design under uncertainty
A tool schema is an API whose caller is probabilistic: descriptions are the documentation the caller actually reads, errors must teach recovery, and schema ambiguity becomes wrong calls at runtime.
Sep 2026 -
Treat MCP servers as supply chain, not plugins
An MCP server is third-party code with prose injected into your context and credentials in its hands — a supply-chain trust decision, not an app-store install.
Sep 2026 -
Voice agents are latency budgets with a language model inside
Voice-agent latency begins with explicit clock boundaries, context-sensitive endpointing, and interruption policies that preserve conversational state.
Sep 2026 -
A Jev tool verdict must bind to the exact action it permits
A semantic verdict cannot mint authority; it must remain bound to the exact action, arguments, principal, tenant, and current policy.
Sep 2026 -
A typed confidence field still needs a local calibration contract
Distribution-derived confidence is not local correctness; consequential thresholds need defined events, held-out evidence, and action-specific fallbacks.
Sep 2026 -
A2A security begins with message authority, not transport encryption
Encrypted transport cannot prove that a delegated task carries current, scoped, replay-resistant authority across every agent hop.
Sep 2026 -
Agent architectures should expose control-state transitions, not framework objects
Agent architectures should expose stable control-state transitions so clients can handle authority, effects, recovery, and completion without framework coupling.
Agentic AI -
Agent capability platforms should expose policy surfaces, not catalog size
Catalog breadth creates options; explicit action, data, identity, revocation, and evidence policies make those options operable.
Agentic AI -
Agent conversation history is shared state, not a transcript
Once messages coordinate agents, their identity, ordering, projections, retention, and replay become state-management concerns.
Agentic AI -
Agent debugging must preserve rejected alternatives, not only chosen actions
Chosen actions explain what happened; rejected tools, plans, and hypotheses reveal whether retrieval, ranking, policy, validation, or budget moved the decision boundary.
Aug 2026 -
Agent frameworks are the new ORMs — the abstraction debt comes due at production
LangGraph, CrewAI, and AutoGen abstract the loop the way ORMs abstract SQL — the demo gets easy, and the debt comes due in production. What frameworks genuinely earn is persistence and contracts; the orchestration vocabulary is where the lock-in lives.
Aug 2026 -
Agent liability follows the authority chain, not the model supply chain
Liability should track who authorised the action, set its scope, could prevent it, and benefited — not default to whichever vendor trained the model.
Aug 2026 -
Agent loops need budget exhaustion as a first-class terminal state
Budget exhaustion is not generic failure—it is a typed terminal state that preserves progress, explains the stop, and keeps agent costs operable.
Sep 2026 -
Agent memory should consolidate by evidence, not recurrence
Repeated observations can share one source or one error; durable memory needs provenance, contradiction handling, authority, and evidence-weighted promotion.
Agentic AI -
Agent observability should record decision deltas, not just spans
Behavior becomes diagnosable when traces preserve changed evidence, candidates, policy, state, and rejected alternatives.
Sep 2026 -
Agent permissions should expire like leases, not persist like roles
Standing roles let agent authority outlive its purpose; task-bound leases make scope, delegation, renewal, and expiry part of the authorization contract.
Aug 2026 -
Agent privacy is a data-flow problem, not a policy document
Privacy is decided at three boundaries — what enters context, what leaves in tool calls, what persists to memory — and enforced with data-engineering artifacts: projections, purpose tags, TTLs, PII pipelines. The policy names intentions; the flows leak facts.
Aug 2026 -
Agent protocols will consolidate the way network protocols did
MCP, A2A, SLIM, and WebMCP are already layering the way networking did; bet on layer boundaries and protocol-independent invariants — identity, delegation chains, audit — not on a single winner.
Aug 2026 -
Agent QA should test recovery paths more than happy-path completion
Agent quality assurance should spend more effort on recovery paths, invariant preservation, and minimal failing trajectories than on repeated happy-path completion.
Sep 2026 -
Agent registries should index revocation before capability
Capability search should begin only after revoked identities, credentials, and trust paths have been removed from the eligible agent set.
Sep 2026 -
Agent routing should buy verified capability, not declared capability
Advertised capability is cheap to state. Route on proven execution history and causal credit, so the pool corrects itself as performance shifts.
Aug 2026 -
Agent state needs migration semantics before it needs faster checkpoints
Checkpoint speed is secondary if new code cannot preserve the meaning of old goals, approvals, pending effects, and control state across releases.
Aug 2026 -
Agent telemetry should preserve decisions, not merely token streams
Useful telemetry preserves candidate choices, rejections, commitment points, and causal context so production runs can become replayable evidence.
Aug 2026 -
Agent tenancy is a state-isolation problem before it is a compute-isolation problem
Multi-tenant agent platforms fail along their state boundaries, not their compute boundaries — session isolation, memory namespaces, and identity scoping decide tenancy correctness before containers do.
Agentic AI -
Agent testing is simulation engineering — record-replay beats mocks
Static mocks test the paths developers predicted; production-grade agent tests replay real trajectories, validate state transitions, and simulate the world the agent actually changes.
Aug 2026 -
Agent usability tests need recovery tasks, not completion ratings
Production trust is measured by whether people can diagnose, repair, and resume failed agent work without creating a second failure.
Agentic AI -
Agent webhooks are distributed transactions wearing HTTP clothes
Reliable callbacks need durable task states, an outbox, explicit delivery transitions, idempotent retries, replay protection, and owned compensation boundaries.
Aug 2026 -
Agent-to-agent trust needs reputation systems, not just protocols
MCP and A2A answer how agents talk, not which agents to believe. Identity proves who an agent is; reputation predicts what it will do — multi-dimensional, topic-gated, telemetry-fed, and designed against Sybils, whitewashing, and the honest-then-malicious pivot.
Aug 2026 -
Agents should not make intervention claims from predictive confidence
Predictive confidence supports forecasts, not claims about what an action will cause; interventions need causal assumptions and identified evidence.
Sep 2026 -
Batch inference needs workflow semantics, not bulk HTTP
Asynchronous inference needs durable identity, explicit lifecycle, deadline-aware scheduling, idempotent reconciliation, and typed terminal outcomes.
Sep 2026 -
Belief-desire-intention models matter only when runtime state is inspectable
BDI becomes architecture only when beliefs, goals, intentions, revisions, and commitments are explicit, versioned runtime objects.
Sep 2026 -
Benchmark scores reward completion; production systems must reward containment
Completion proves capability; containment proves the capability belongs in production by scoring boundary adherence, recovery quality, side effects, and bounded resources.
Aug 2026 -
Browser agents should prefer accessibility trees before pixels
Start with declared interface semantics, then escalate to visual reasoning only when the interface withholds usable structure.
Agentic AI -
Cache keys are behavior contracts, not performance hints
A cache key defines which requests may share behavior; meaning, authority, evidence, and effect boundaries must all participate in identity.
Agentic AI -
Cancellation is a protocol, not an exception path
Graceful stop must be formalized like success: a state machine with acknowledgement, bounded unwind, and a traceable record — not a best-effort signal.
Aug 2026 -
Cheap routing decisions must survive the cost of fallback
Routing economics belong on the completed-task ledger, including retries, repair, fallback, latency, and human review.
Sep 2026 -
Code sandboxes must meter burstiness, not average utilization
Agent-generated code alternates between silence and sharp resource peaks; tool-call-scale metering and graduated controls contain contention that averages hide.
Sep 2026 -
Code-as-action beats JSON tool calls once tasks compose
CodeAct showed up to 20% higher success and ~2x fewer steps: code composes tools with loops and variables where JSON spends a round-trip per call. The advantage has a competence threshold and a security bill — match the action space to the task's compositionality.
Aug 2026 -
Coding agents should retrieve dependency boundaries before nearby text
Repository context should begin with interfaces, dependencies, and callers that constrain an edit—not whichever text happens to be nearby.
Sep 2026 -
Cold starts are part of agent correctness, not startup optimization
Readiness must prove identity, policy, schema, state, and capability compatibility before an agent accepts work; fast partial initialization is still incorrect.
Aug 2026 -
Composite decision scores need an explicit policy for disagreement
Weighted scores hide policy unless precedence, uncertainty, and non-compensable concerns remain explicit.
Sep 2026 -
Computer-use agents are the integration of last resort — design for when APIs don't exist
Pixel-in, action-out is the most expensive and most fragile way to integrate with software — treat GUI automation as the fallback for when APIs don't exist, and engineer the fallback like the hazard it is.
Aug 2026 -
Computer-use evals should verify resulting state, not clicks
GUI-agent evaluation should inspect authoritative application state, preserved invariants, and duplicate effects rather than reward plausible cursor choreography.
Sep 2026 -
Context caches need invalidation contracts, not hit-rate targets
A fast hit on stale policy, evidence, schemas, or session state is a behavioral defect; cache identity needs owned versions, scope, freshness, and fallbacks.
Agentic AI -
Context-compression thresholds are reliability budgets
Compression timing allocates protected evidence, usable headroom, and the amount of information loss a long-horizon run may safely spend.
Sep 2026 -
Conversation management needs state ownership, not longer transcripts
Continuity, ordering, branching, retention, and recovery require an explicit state owner; a larger context window only postpones the decision.
Agentic AI -
Cost regression gates belong in CI next to eval gates
Quality and cost are two dimensions of the same production contract — a change that makes an agent cheaper but fails more in production isn’t cheaper.
Aug 2026 -
Cross-lab safety evals need a shared failure ontology
Comparative safety claims need shared definitions for failures, opportunities, operating conditions, containment, and grader uncertainty.
Sep 2026 -
Cross-system harnesses need a canonical event model, not adapter sprawl
One owned event model should preserve identity, lifecycle, effects, errors, and evidence while protocol adapters declare semantic loss.
Sep 2026 -
Debate and ensembles buy reliability with tokens — price the trade explicitly
Sampling, mixture-of-agents, and debate are redundancy priced in tokens; write down the exchange rate, because a verifier or uncertainty-routed escalation often buys the same reliability cheaper.
Aug 2026 -
Decision-tool comparisons should name the control each tool owns
Classifiers, routers, guardrails, policy engines, and agent hosts belong in different control layers, not one universal leaderboard.
Sep 2026 -
Deploy agents where their data is allowed to exist, not where inference is cheapest
Resolve residency and sovereignty before cost, then enforce the result through regional queues, geo-locked inference, local state, egress policy, and provenance.
Aug 2026 -
Drift detection is the production eval — offline suites only catch what you predicted
Prompt drift, provider drift behind unchanged model IDs, slow quality decay — a frozen suite is structurally blind to all of it. Control charts, judge sampling, and gated rollouts are the eval that runs on truth; invert the eval budget toward monitoring.
Aug 2026 -
Emergent personas are release regressions, not stylistic quirks
Cross-surface behavioural changes after narrow training updates belong in release gates, shadow evaluations, and rollback plans.
Agentic AI -
Every agent extension spends governance before it creates capability
A new skill, plugin, hook, or MCP server adds a new principal, instruction source, dependency, and failure mode — extensibility is a transfer from a finite governance budget.
Aug 2026 -
Fine-tuning is a data-asset decision, not a model decision
LoRA and QLoRA made the tuning run cheap and disposable. The dataset — curated instructions, preference labels, filtered trajectories — is the asset that survives base-model churn, and it deserves data-engineering discipline.
Aug 2026 -
Guardrails are a product surface, not a safety bolt-on
Refusal copy, false-positive rates, gate frequency, and guard latency are what users actually experience at the boundary — the soft guardrail layer is product design, resting on a hard deterministic floor that isn't.
Aug 2026 -
Guardrails belong at the action-decision layer, not the output filter
Safe-looking prose cannot contain unsafe effects; production guardrails must evaluate typed action proposals before tools execute.
Agentic AI -
Handoffs are the riskiest primitive in multi-agent design
A handoff transfers ownership — context, authority, and the user’s intent — across an agent boundary in one lossy move, and SDKs make it look like a function call. Scoped calls by default; ownership transfer as the engineered exception.
Aug 2026 -
Human feedback is a control loop — unmeasured disagreement becomes hidden bias
When annotators disagree, that disagreement is information about ambiguity, competing values, or uneven expertise — collapsing it into one target hides the conflict inside the reward model.
Aug 2026 -
Human oversight works only when confidence changes the handoff
Confidence is useful only when it changes who acts, who reviews, or whether an agent must stop.
Sep 2026 -
Jev belongs in the decision layer, not in place of the agent framework
A structured-decision model can replace one uncertain judgment; orchestration, recovery, policy, and execution remain application responsibilities.
Sep 2026 -
JSON repair must never invent authorization-bearing values
Repair may restore syntax; it must not manufacture approvals, targets, operations, or any other value that can authorize an effect.
Sep 2026 -
Loop completion needs user-confirmed terminal conditions, not model self-report
A model may propose completion, but the harness must verify authoritative state and route qualitative acceptance to the user through typed terminal outcomes.
Agentic AI -
Memory retrieval needs explicit query classes before a universal ranker
Recency, identity, procedural, semantic, and investigative questions need declared plans and hard predicates before their eligible evidence is ranked.
Agentic AI -
Memory scope is a schema axis, not a namespace prefix
User, task, agent, team, and tenant scope change memory ownership, lifecycle, authority, and retrieval precedence; a key prefix cannot carry that contract.
Aug 2026 -
Multi-agent debate has a coordination budget, not free accuracy
Debate improves decisions only while independent evidence exceeds the coordination, drift, latency, and verification costs added by more agents and rounds.
Sep 2026 -
Multi-agent economics fail without causal credit assignment
You can’t price, optimize, or scale a fleet if you don’t know which agent produced each outcome. Attribution traces the tool-call graph, not the invoice.
Aug 2026 -
Multi-agent systems need negotiated backpressure, not faster chat
The production problem is whether the receiver can declare what it can accept, the sender can slow down without losing intent, and both can preserve a truthful account of what was promised.
Aug 2026 -
Multi-agent topology is a cost model before it is an orchestration diagram
Star, chain, tree, and mesh topologies price context duplication, critical-path depth, verification, attribution, and failure propagation differently.
Sep 2026 -
Multilingual agents need per-language action evals, not translated prompts
Language support becomes a production claim only when native scenarios verify tool choice, authority, recovery, and resulting state.
Sep 2026 -
On-device models redraw the agent privacy and latency boundary
1–7B models, mature quantization, and sub-50ms local tokens make placement the real design question: partition agent functions by latency class and data sensitivity — and let memory bandwidth, not TOPS, decide what fits.
Aug 2026 -
Open agent ecosystems make identity lifecycle more important than discovery
Discovery lists agents; issuance, lineage, delegation, rotation, revocation, and retirement determine whether open populations remain governable.
Sep 2026 -
Output watermarks are provenance signals, not proof of authorship
Watermark detection is probabilistic evidence whose meaning depends on length, entropy, editing, key custody, and false-positive cost—not an authorship verdict.
Sep 2026 -
Parallel tool calls need effect isolation before they need concurrency
Fan-out is safe only when effect classes, dependencies, partial failures, retries, and commit states remain explicit.
Agentic AI -
Privacy budgets are runtime resources, not compliance numbers
Differential-privacy budgets must be allocated, composed, metered, exhausted, and renewed by the runtime—not recorded as static compliance values.
Sep 2026 -
Procedural skills should carry preconditions and expiry, not just instructions
Reusable agent procedures need applicability checks, ownership, versions, and expiry so correct instructions are not applied to the wrong world.
Sep 2026 -
Prompt governance needs release semantics, not a template library
Immutable effective configurations, evaluation gates, protected promotion, traceable rollout, compatibility, and rollback turn prompt edits into governed releases.
Sep 2026 -
Prompt optimization is compilation — prompts deserve version control
OPRO, DSPy, and ProTeGi turn intent plus evals into prompt text. Once prompts are compiled artifacts, hand-editing them is patching a binary — version the sources, pin the model, and gate changes with regression suites.
Aug 2026 -
Read-only knowledge should be an MCP resource, not another executable tool
Stable knowledge belongs in application-driven resources; executable tools should remain reserved for computation and effects.
Sep 2026 -
Reasoning models change the economics of planning, not the need for it
Test-time compute made planning quality purchasable per query — but self-verification still fails, external verifiers still decide what's true, and planning discipline survives the new price list.
Aug 2026 -
Red teaming agents is an engineering practice, not an audit event
An agent's attack surface shifts with every model, prompt, and tool change — red teaming has to run continuously in CI with utility-under-attack metrics, not as a point-in-time audit.
Aug 2026 -
Reflection without explicit stopping criteria is recursive cost, not self-improvement
Critique and revision become a learning mechanism only when quality deltas, resource limits, and escalation boundaries say whether another iteration is justified.
Aug 2026 -
Reproducible agents require controlled boundaries, not deterministic models
You can’t make an LLM deterministic; you can only constrain the space it samples from. Reproducibility is a boundary problem — pinned versions, frozen snapshots, ordered execution — not a seed value.
Aug 2026 -
Response APIs make conversation state an infrastructure dependency
Hosted continuity, reasoning, tools, and background work need explicit ownership, retention, and recovery boundaries.
Agentic AI -
RL-trained harnesses need policy-version boundaries around tool use
RL-trained harnesses should bind each episode to the model, tool surface, permissions, and verifier that produced its actions, traces, and rewards.
Agentic AI -
Runtime contracts are the safety case for probabilistic agents
When outputs are stochastic, verification moves from the model to the runtime: design-by-contract at every boundary turns each call into a deterministic check.
Aug 2026 -
Sabotage evals must create opportunities, not ask about intentions
An agent can reason perfectly about harm and still execute it. Measure what the architecture permits under pressure, not what the model can explain.
Aug 2026 -
Security agents should earn containment authority separately from detection accuracy
Detection quality does not grant containment authority; judgment, permission, execution, and recovery need separate contracts.
Sep 2026 -
Self-modifying agents must not control their own acceptance tests
A self-modifying agent may propose a better implementation; it must not control the evidence or verdict that permits the change to advance.
Sep 2026 -
Self-play gains need outside opponents before production claims
A policy has not escaped its curriculum until independent opponents, runs, and partners test it outside the training lineage.
Sep 2026 -
Semantic caching is the only cache that can lie to you
A classic cache returns the right answer or nothing. A semantic cache can return a confident wrong answer at cache speed — so operate it like a retrieval model with thresholds and evals, not like a Redis instance.
Aug 2026 -
Serialising agent state is a compatibility contract, not a persistence detail
Serialised state is a message between versions and subsystems—not merely bytes at rest—so ownership, schema compatibility, migration, and rejection must be explicit.
Sep 2026 -
Shadow evals should precede every model promotion
Paired evaluation on eligible production traffic exposes behavioral, trajectory, cost, latency, and segment regressions before users carry the release risk.
Sep 2026 -
Social-agent simulations need population interventions, not plausible conversations
Plausible dialogue cannot explain population behaviour; social-agent simulations need interventions across topology, visibility, identity, and incentives.
Agentic AI -
Stateless MCP shifts state responsibility to clients, not away
Stateless MCP removes protocol sessions by making continuity explicit—clients must carry metadata, continuation tokens, handles, and recovery responsibility.
Agentic AI -
Static IAM cannot express delegated agent authority
Human authorization is decided once; agent authority emerges continuously through tool calls. Capability tokens with scoped delegation chains replace static roles.
Aug 2026 -
Streaming isn’t UX polish — it’s the agent runtime’s backpressure design
Streaming becomes runtime architecture when it regulates producers, consumers, concurrency, cancellation, and failure — not when it merely animates tokens.
Aug 2026 -
Structured-output evals must test recovery, not just valid JSON
Reliable structured output requires semantic checks, typed failure classes, bounded repair, cost-aware retry, truncation detection, and safe terminal outcomes.
Sep 2026 -
Task decomposition is where agent plans go to die
When a long-horizon run fails, the execution trace gets the blame but the decomposition committed the crime: granularity, dependency structure, and goal propagation are schema decisions whose errors compound downstream.
Aug 2026 -
The EU AI Act turns agent governance into a design input
Articles 9, 12, and 14 read as an architecture spec: tamper-evident logging, intervention surfaces, iterative risk management. The Act's agentic gaps make conservative design the rational default — and Annex III enforcement began this month.
Aug 2026 -
The LLM gateway is a control point, not plumbing
Routing, fallbacks, budgets, tenancy, residency, model pinning — the gateway is the one chokepoint every token crosses, which makes it where cost, reliability, and compliance policy actually execute. Treat its config as versioned policy code.
Aug 2026 -
The schedulable unit of agent work is the effect, not the pod
A pod’s lifecycle maps to wall-clock; an agent’s value maps to outcomes. Schedule for verified effects, not occupied compute.
Aug 2026 -
Tool errors must be typed for the agent, not logged for the operator
Typed failures tell the agent whether to repair, retry, choose another tool, compensate for an effect, or stop and escalate.
Sep 2026 -
Tool retrieval should optimize capability coverage, not nearest-neighbour similarity
A useful tool set covers the operations a task requires; similarity is only candidate generation, not the acceptance criterion.
Sep 2026 -
Tool schemas are security boundaries, not developer documentation
A tool schema defines which actions a probabilistic caller can express; ambiguity and permissive defaults become unsafe authority before policy runs.
Sep 2026 -
Tool schemas should be pinned like dependencies, not rediscovered on reconnect
Live tool discovery silently changes executable authority; pin normalized manifests, verify reconnects, and promote schema changes through a governed release boundary.
Agentic AI -
Tool security starts by removing capabilities, not detecting misuse
Task-scoped read, execute, and write authority should make dangerous effects impossible before probabilistic detection is asked to recognize misuse.
Agentic AI -
World models earn their cost only when environment drift defeats prompt updates
A predictive environment model is justified when changing action dynamics defeat fresh retrieval, explicit rules, and prompt updates.
Sep 2026
Data platform & strategy
59 notes
-
AI-ready data is modeled data
There is no shortcut through embeddings. Data that can't support analytics semantics can't support an agent's decisions either.
Aug 2026 -
Data contracts are the API layer of the data platform
Schema change without negotiation is the root cause of most pipeline breakage and silent AI degradation; contracts turn data interfaces into APIs with semantics.
Aug 2026 -
Data products fail at lifecycle, not design
Launch must begin a governed operating loop of ownership, outcome evidence, versioning, and retirement—not end a delivery project.
Sep 2026 -
Streaming is the default temperature of agent-era data
Agents act now, on current state — batch-refreshed data that was fine for human dashboards becomes an error source for autonomous action. Streams, real-time OLAP, and activation paths shift from luxury to default.
Sep 2026 -
The AI flywheel is a data architecture
A real flywheel names its product, operations, or ecosystem loop and gives every evidence handoff an observable owner.
Data platform & strategy -
The lakehouse is the natural substrate for agent memory
Agent memory needs cheap append, time travel, schema evolution, ACID merge, and multi-engine access — exactly what open table formats and the lakehouse already deliver.
Aug 2026 -
The semantic layer is what your LLM actually queries
Text-to-SQL against raw schemas fails on semantics, not syntax; a semantic layer makes NL2DB reliable by moving meaning from the prompt into the platform.
Aug 2026 -
Agentic analytics needs semantic guardrails more than conversational polish
Fluent answers cannot rescue the wrong metric, grain, join path, time window, or access policy; govern meaning before polishing the conversation.
Aug 2026 -
AI data quality is a distribution contract, not a null-check suite
Structurally valid records can still break an AI system when population mix, coverage, duplication, or semantic distributions drift beyond the behaviour it was built to preserve.
Aug 2026 -
Approximate-nearest-neighbour recall is a product SLO
ANN miss rates are user-visible reliability failures, so recall must be sliced, budgeted, and release-gated.
Sep 2026 -
Arrow and Parquet are the lingua franca of the AI data stack
The composability of the modern data stack is a byte-layout agreement: Arrow in memory, Parquet on disk, zero-copy at every boundary. Standardize on the format pair and the engines become interchangeable parts.
Aug 2026 -
Bi-temporal modeling is table stakes for agent-era data platforms
Every agent asks “what’s true now?” and “what did we believe when we decided?” — and single-timeline platforms can only answer the first. Two time axes are now the auditable default.
Aug 2026 -
CDC correctness is identity plus ordering, not connector throughput
A connector that moves 100k rows but loses record identity or event order produces corrupt state. Correctness lives in identity guarantees and sequence integrity.
Aug 2026 -
Choose OLAP engines by query shape and operating model, not benchmark rank
Benchmarks run synthetic workloads on clean data. Match the engine to your actual query patterns and the team that operates it.
Aug 2026 -
Classifier label changes are downstream data-contract changes
A stable JSON shape can still break consumers when category meaning, membership, or routing semantics change.
Sep 2026 -
Continual model updates need knowledge-change ledgers
Model updates need a queryable history of intended corrections, supporting evidence, retention tests, and recovery choices.
Sep 2026 -
Data products need retirement metrics before portfolios become data swamps
Adoption, decision value, operating cost, and successor readiness must justify continuation; catalog size without subtraction is governed sprawl.
Sep 2026 -
Data provenance is rollback infrastructure for AI behavior
When a model misbehaves, the only thing you can fix retroactively is its training data. Provenance chains make targeted rollback possible instead of full retraining.
Aug 2026 -
Data rights belong in lineage, not legal footnotes
Licences, consent, purpose, residency, and deletion obligations must travel with data so pipelines can enforce them and calculate the blast radius of change.
Aug 2026 -
Data Vault vs dimensional: model for rate of change, not fashion
The vault-vs-star argument is resolved by measurement, not loyalty: count your sources, measure their churn, check your audit obligations. The vault absorbs change in Silver; the star provides meaning in Gold — a pipeline, not a rivalry.
Aug 2026 -
Data visualizations are semantic interfaces, not chart outputs
Every scale, encoding, label, hierarchy, and state changes the claim a reader receives and the decision the interface can safely support.
Sep 2026 -
Dataset previews are governance interfaces, not conveniences
A preview should expose contracts, provenance, policy, and representative risk before consumption becomes dependency.
Sep 2026 -
Decision-model replay needs versioned state, questions, and labels
Replay can explain model change only when evidence, question definitions, labels, and consuming policy are versioned together.
Sep 2026 -
Deduplicating training data is capacity allocation, not housekeeping
Each dedup pass spends FLOPs to save tokens, at decreasing margin. Where you stop is a resource decision between volume and fidelity, not a cleanup task.
Aug 2026 -
Differential-privacy budgets belong in data-product contracts
Privacy promises need protected units, composition scope, utility boundaries, lineage, and exhaustion behaviour.
Sep 2026 -
Differentially private training needs accounting lineage across runs
Privacy spend composes across retries, checkpoints, overlapping datasets, and derived releases—not only the successful job a team kept.
Sep 2026 -
Document extraction is the unglamorous bottleneck of enterprise RAG
Enterprise RAG fails at ingestion, not retrieval — the hard problem is extracting structured content from PDFs, DOCX, and code with semantic fidelity.
Aug 2026 -
Ecosystem convenience spends portability through hidden data conversions
Convenient dataset abstractions stay portable only when schemas, lineage, streaming semantics, and rebuild paths remain explicit.
Sep 2026 -
Embedded stores need explicit export and repair paths
Embedding a database removes an operator, not the need for consistent export, evidence-preserving repair, and verified restoration.
Sep 2026 -
Embedded vector data needs exportable formats before local speed
Local retrieval is production-ready only when records, schemas, embedding provenance, and rebuild instructions can leave the engine intact.
Sep 2026 -
Embedded vector stores trade network operations for lifecycle discipline
Removing a service boundary moves schema, index freshness, compaction, recovery, tenancy, and measured recall into the application lifecycle.
Sep 2026 -
Embedding migrations need dual-read cutovers, not one-shot reindexing
Shadow indexes and dual reads turn vector-space changes into measured, reversible production migrations.
Sep 2026 -
Entity identity is the first data-model decision, not a surrogate-key afterthought
A generated key distinguishes rows but cannot decide which real-world thing persists, when two records are the same, or how identity survives change.
Sep 2026 -
Episodic memory retrieval needs SQL and semantics, not embeddings alone
Episodic memory retrieval needs deterministic SQL for identity, time, and scope before semantic ranking can help with meaning.
Sep 2026 -
Exact identifiers deserve lexical retrieval before semantic fusion
Identifiers, codes, and quoted phrases carry exact intent, so lexical retrieval should establish candidates before semantic fusion broadens recall.
Sep 2026 -
Feature pipelines need point-in-time contracts before reusable stores
Reusable features need explicit event time, availability, correction, and offline-online semantics before a store can distribute them safely.
Sep 2026 -
Federated learning needs participant provenance before averaging
Federated learning needs participant eligibility, training context, and contribution history before averaging.
Sep 2026 -
Forecasting foundation models still need local baselines
A pretrained forecaster earns production trust only when it beats a local control on the series, horizon, calibration, and drift behavior that matter.
Sep 2026 -
Hub-and-spoke activation centralizes data authority as well as distribution
A warehouse hub reduces pipeline sprawl while making one governed model authoritative across every operational spoke.
Sep 2026 -
Knowledge graphs earn complexity when relationships carry operational meaning
A graph earns its cost when typed paths change permissions, impact, provenance, inference, or action—not merely when connected data looks useful.
Sep 2026 -
Metadata is the control plane of the agentic data stack
When agents become first-class data consumers, metadata stops being documentation: what an agent can discover, query, and be permitted to do is executed through catalogs, contracts, lineage, and policy — not described by them.
Aug 2026 -
Normal forms are change-control tools, not academic purity
Normalization localizes change, exposes dependency ownership, and prevents one business fact from acquiring conflicting stored versions.
Sep 2026 -
Platform consolidation buys governance by spending portability
Consolidation makes governance coherent by coupling policy, semantics, execution, and operations to one platform — open storage lowers the price of exit without making exit free.
Aug 2026 -
Poisoning resistance begins at ingestion, not alignment
Provenance, quarantine, staged promotion, and reversible lineage make poisoning a governed data-path problem before alignment begins.
Aug 2026 -
Preference datasets need disagreement topology, not majority labels
Preference data should preserve differences across criteria, evidence, response segments, and judges.
Sep 2026 -
Prompt pools are versioned training data, not reusable text snippets
Adaptive prompt pools need provenance, joint router versioning, retention evaluations, and reproducible learning history.
Sep 2026 -
RAG evaluation should separate retrieval failure from answer failure
End-to-end scores hide whether evidence was missing or merely mishandled; evaluate retrieval and generation as separate failure surfaces.
Sep 2026 -
Reverse ETL is an operational write path, not a marketing sync
Writing warehouse decisions into business systems requires identity, consent, idempotency, reconciliation, and recovery discipline.
Sep 2026 -
Schema evolution is producer-consumer negotiation, not parser permissiveness
Safe change needs compatibility claims, impact analysis, migration windows, and consumer acknowledgement — not merely a writer that accepts new columns.
Sep 2026 -
Storage-compute separation turns engine choice into a reversible decision
Storage-compute separation makes engine choice reversible only when open table state, catalog semantics, and maintenance are portable too.
Sep 2026 -
Streaming storage should be selected by replay and compaction economics, not latency alone
The retained path determines the cost of replay, compaction, recovery, and onboarding consumers; live-path latency is only one lifecycle constraint.
Sep 2026 -
Tabular foundation models must beat gradient-boosted local baselines
Foundation models earn a production role only by beating credible, leakage-safe boosted-tree baselines on the local decision problem.
Sep 2026 -
Tokenization changes are data migrations, not model upgrades
A tokenizer change alters derived representations, chunk boundaries, context budgets, and historical usage units, so it needs versioning and rollback.
Sep 2026 -
Training scale should be governed by data yield, not accelerator count
Approve compute in stages and require each tranche to earn the next through measured marginal learning.
Sep 2026 -
Training-data mixtures are product priorities expressed as weights
Domain weights allocate finite model capacity among promised behaviours, users, and risks—the product roadmap made numerical.
Aug 2026 -
Vector multi-tenancy fails at metadata isolation before retrieval relevance
A relevant result from the wrong tenant is a security incident; tenant identity must be an enforced index contract, not a caller-supplied filter.
Sep 2026 -
Vector search is an indexing decision, not a database purchase
Embedding search is an index type with HNSW/IVF trade-offs, filtering, and tenancy concerns — increasingly living inside databases you already run. Evaluate it like an index, not a category purchase.
Aug 2026 -
Vision-language datasets need interaction coverage, not image volume
Dataset quality depends on coverage of instructions, evidence demands, reasoning forms, and failure boundaries—not raw image count.
Sep 2026 -
Your observability stack is a data platform wearing a dashboard
Agent observability is not a dashboard problem. It is an event data platform: trace schemas, cost records, eval scores, lineage, retention, and query shape decide what production behavior can be explained.
Aug 2026
Architecture & systems
23 notes
-
Normalization is a spectrum, not a rule
Third normal form is a starting position, not a destination. Where you sit on the normalization spectrum is a workload decision.
Jun 2026 -
Consistency models are product decisions
Consistency design separates transaction anomalies, replica freshness, tenant isolation, and deterministic replay into explicit product contracts.
Sep 2026 -
Alternative sequence models should be chosen by state-update economics, not context length
Compare bytes moved, update latency, memory tiers, and retained evidence on the workload that will run.
Sep 2026 -
Consensus protocols buy agreement, not application correctness
A committed replicated history can still contain the wrong command; business invariants, authority, freshness, and side effects remain application work.
Sep 2026 -
Distributed compute abstractions leak at checkpoint and object-store boundaries
Remote functions feel local until state, object movement, checkpoint durability, and causal replay cross process and storage boundaries.
Sep 2026 -
Durable workflow histories are executable recovery contracts
Recorded events must constrain replay, effect handling, code compatibility, and the next valid action after failure.
Sep 2026 -
Durable workflows should make interruption a state transition
Pause, cancellation, and replanning must be recorded, replayable transitions with explicit step state and deduplication boundaries.
Sep 2026 -
Event logs need replay ownership before event sourcing
An append-only log becomes a recovery architecture only when someone owns historical compatibility, reconstruction, and replay-safe effects.
Sep 2026 -
Event sourcing earns its complexity only when replay is the feature
Most CRUD apps never replay anything and pay the event-sourcing complexity tax forever; agent systems are the rare case where replay IS the product — debugging, audit, and resume all consume the log.
Aug 2026 -
Harness-compute separation makes model runtimes replaceable
Keep reasoning, authority, workspace declaration, and recoverable state in the trusted harness so execution compute can be lost or replaced.
Architecture & systems Sep 2026 -
Idempotency is the cheapest reliability you will ever buy
Retries are the universal failure response, and retries without idempotency convert failures into duplicates; designing safely repeatable operations is decades-old discipline agentic tool execution makes urgent again.
Aug 2026 -
Local-first AI turns file watching into a consistency protocol
Watch events are low-latency evidence inside a journaled protocol of generations, tombstones, reconciliation, and visible freshness.
Aug 2026 -
Long context shifts retrieval errors into attention allocation
Long windows move failure from document selection into attention allocation, position, and cache management.
Sep 2026 -
Postgres is the default answer — agent state included
PostgreSQL has become the default platform for structured state — and agent state management is no exception, from conversation history to durable checkpoints.
Aug 2026 -
Protocol adapters are permanent architecture, not migration scaffolding
Adapters own semantic loss, lifecycle translation, security policy, observability, and conformance across heterogeneous agent edges.
Aug 2026 -
Recovery architecture begins with an explicit consistency-anomaly budget
Name the stale reads, duplicates, lost updates, and partial effects recovery may expose—and bound them by audience and time.
Sep 2026 -
Recovery plans start with state ownership, not restart commands
Recovery begins by identifying authoritative state, committed effects, and reconstruction paths; restarting is only one mechanism.
Sep 2026 -
Routing tables are policy interfaces, not infrastructure configuration
Target selection allocates capability, cost, reliability, and authority, so routes need versioned policy and decision evidence.
Sep 2026 -
Rust error handling is an architecture for recovery, not syntax
Rust error types define which failures callers can distinguish, which causes survive, and where retry, containment, or escalation is allowed.
Architecture & systems Sep 2026 -
Service meshes taught us what agent meshes will relearn
Service meshes solved zero-trust networking, traffic management, and observability for microservices — agent meshes will relearn these lessons for multi-agent systems.
Aug 2026 -
Stateful concurrency belongs in actors before containers
Actors own logical identity, ordering, supervision, and passivation; containers remain the physical isolation and placement layer underneath.
Aug 2026 -
Supervision trees contain failure better than global retry loops
Recovery should follow dependency boundaries: supervision trees preserve healthy work, cap restart storms, and escalate only the state that can no longer be trusted.
Aug 2026 -
The operating system is the right mental model for agent runtimes
Context window as RAM, tools as device drivers, permissions as user vs kernel space, the loop as a scheduler with budgets — the analogy imports fifty years of resource-management discipline, right up until the kernel turns out to be stochastic.
Aug 2026
Product & strategy
13 notes
-
Specs as forcing functions
A spec's primary value isn't documentation — it's forcing decisions to be made while they're still cheap.
Jun 2026 -
Agent marketplaces will be won by trust infrastructure, not catalog size
Two-sided agent marketplaces face the same trust problem as early e-commerce — and the winners will be those who solve authorization, authenticity, and accountability first.
Aug 2026 -
Agent moats come from embedded workflow data, not model exclusivity
Exclusive capability decays; embedded workflows compound verified outcome data into durable advantage.
Sep 2026 -
Agent product defensibility decays unless workflow data compounds
A launch advantage becomes a moat only when verified workflow outcomes improve the product faster than competitors can copy it.
Sep 2026 -
Agent products should own the exception queue before the happy path
The exception queue exposes domain judgment, creates proprietary feedback, and lets an agent earn autonomy before routine automation becomes the goal.
Aug 2026 -
Agent products should publish operating envelopes, not capability lists
State the tasks, environments, authority, horizons, budgets, and supervision under which capability remains dependable — then widen that envelope with evidence.
Aug 2026 -
AI billing needs causal usage attribution before price innovation
Before pricing tokens, tasks, agents, or outcomes, preserve the causal chain from customer intent through runs, cost, and billable effect.
Sep 2026 -
AI moats require rights to learn from workflow data
Production evidence compounds only when contracts permit it to improve memory, evaluation, policy, retrieval, or models at an agreed scope.
Sep 2026 -
AI pricing must expose workload variance before margins
Pricing must expose the token mix, reasoning cost, and delivery constraints hidden beneath each customer-facing unit.
Sep 2026 -
AI-native businesses price outcomes, not seats
Seat pricing assumes value scales with humans logged in — exactly the assumption AI products break. Usage and outcome pricing demand metering infrastructure and margin discipline most SaaS companies don't have.
Aug 2026 -
Enterprise AI buying is a trust purchase — sell governance, not magic
The demo sells capability; procurement buys accountability. Audit trails, permission models, and certifications decide enterprise AI deals — the graduated trust ladder is the real product, so put the audit trail in the demo.
Aug 2026 -
Vertical agents should encode regulatory workflow, not domain vocabulary
A vertical product is policy, memory, tools, verification, environment, and exception ownership assembled around a regulated job—not a vocabulary pack.
Sep 2026 -
Vertical agents win on domain constraints, not model quality
Vertical agent products defend on encoded domain constraints, workflow integration, compliance postures, and proprietary feedback data — not on having a better model. The constraint set is the product.
Aug 2026
UI/UX & design
13 notes
-
Trust is a design material in AI interfaces
The goal is calibrated trust, not maximum trust. Approval UX is a security property, and every permission dialog is spending a budget.
Aug 2026 -
Accessibility is system correctness, not interface polish
Accessibility is a behavioral contract across semantics, state, interaction, content, and recovery—not a final visual compliance pass.
Sep 2026 -
Agent interfaces should show recovery state before reasoning text
Show effects, uncertainty, checkpoints, and safe next actions before asking users to interpret model reasoning.
Sep 2026 -
Approval UX should communicate blast radius, not ask for generic consent
Meaningful approval names the target, action, maximum effect, reversibility, and lifetime of authority instead of asking whether the user generically trusts the agent.
Aug 2026 -
Design systems are constraint systems — and generative UI makes that literal
A design system was always a constraint system pretending to be a component library. Generative UI makes it literal: machine-readable tokens and rules become the grammar the model composes in — and the ablations price the constraint set at ~100 ELO.
Aug 2026 -
Design tokens are the contract that makes AI-generated UI governable
When AI generates UI, design tokens become the governance layer — the typed contract between creative intent and model output that keeps interfaces consistent at scale.
UI/UX -
Form errors should repair intent, not merely reject input
A useful validation failure preserves work, identifies the mismatch, and makes the next correct action obvious.
Sep 2026 -
Generative UI moves layout into the model's output space
When the model emits interface rather than paragraphs, design systems become grammars the model composes in — the designer's job shifts from drawing screens to defining the constraint system the generation must obey.
Aug 2026 -
Governance controls that leave the admin console do not exist for operators
Policy becomes operational only when scope, evidence, approvals, staleness, and rollback are visible where administrators work.
Sep 2026 -
Heuristics vs. aesthetics: where they collide
When visual elegance and usability heuristics pull in opposite directions, which one yields — and why.
Jun 2026 -
In AI interfaces, recoverability beats explainability
An AI interface earns trust by making mistakes cheap to detect, interrupt, correct, and undo — not by narrating an opaque model more fluently.
Aug 2026 -
Terminal agent UX should expose resumability before animation
Named sessions, blocked state, retained boundaries, and reattachment earn trust across absence; motion only proves the renderer is alive.
UI/UX & design Sep 2026 -
Visual hierarchy should expose control priority before brand personality
Operational interfaces should make control state, consequence, evidence, and reversibility visually dominant before they express brand personality.
UI/UX & design Sep 2026
Platform engineering
12 notes
-
Agent CI/CD must version behavior, policy, and state as one release
The deployable identity of an agent spans code, models, prompts, tools, policy, evals, and persisted-state meaning.
Sep 2026 -
Agents are just workloads — Kubernetes discipline applies
Quotas, gang scheduling, isolation profiles, readiness probes, scale-to-zero — every "agent infra" requirement has a mature platform primitive. The substrate is solved; spend the novel engineering on the semantic layer, where agents genuinely are new.
Aug 2026 -
Autoscale agent fleets on queue age and checkpoint cost, not CPU
Agent capacity should follow work debt: queue age, service deadlines, dependency saturation, and the cost of interrupting stateful work.
Aug 2026 -
Desktop AI must treat the renderer as untrusted
Model output and page content stay outside host authority; validated, scoped IPC is the only bridge.
Sep 2026 -
GitOps is the deployment model agents were waiting for
Agents change fast and need audit, rollback, and drift detection more than any workload before them — declarative desired state in git with reconciliation is precisely that discipline.
Aug 2026 -
Health checks should measure recoverable user outcomes
Green should mean a representative outcome can complete—or fail into a preserved, containable, recoverable state.
Sep 2026 -
Kubernetes AI platforms need workload classes before autoscaling
Define latency, state, accelerator, interruption, and tenancy classes before tuning autoscaling signals.
Sep 2026 -
Local Kubernetes environments should optimize feedback fidelity, not production imitation
A local cluster earns its cost by preserving production behaviors that affect decisions while shortening the path to trustworthy feedback.
Platform engineering Sep 2026 -
Local Kubernetes parity starts with failure modes, not manifests
A local environment is faithful when it can provoke, expose, and recover from the production risks the test claims to cover.
Platform engineering Sep 2026 -
Policy-as-code is how enterprises say yes to agents
The alternative to policy-as-code is policy-as-meetings — and agents move too fast for meetings. Evaluable, versioned, auditable policy converts 'too risky' into 'risky actions denied by policy X', which is a yes.
Aug 2026 -
Rate limits are fairness policy for agent fleets, not API protection
Shared fleets need explicit allocation across tokens, time, tools, and tenants; global provider limits cannot prevent noisy-neighbor starvation.
Aug 2026 -
Secrets discipline is a prerequisite for agent fleets, not an afterthought
Agent fleets multiply the secrets surface area — dynamic credentials, per-tenant isolation, and automatic rotation aren't optional security features, they're the prerequisite that makes fleets operable.
Platform Engineering
Teaching & training
13 notes
-
How 32,000+ notes become a curriculum
Atomic notes, an explicit taxonomy, and hub navigation turn accumulation into teachable structure — curriculum design becomes a query over a knowledge graph instead of a blank outline.
Aug 2026 -
Progressive complexity in training design
Sequencing learning so each layer is load-bearing for the next — and why most training fails at layer two.
Jun 2026 -
Supervise agents the way you grow junior engineers
Autonomy is granted in scopes and earned by track record — read-only before write, review-everything before spot-check, narrow domain before broad. Skipping rungs fails the same way for agents and juniors alike.
Aug 2026 -
AI training fails when it teaches tools instead of judgment
Tool knowledge depreciates at the vendor's release cadence; judgment — when to delegate, how to specify, how to verify, when to distrust — compounds. Teach the tool as the vehicle, not the destination, and preserve the work that keeps people qualified to supervise.
Aug 2026 -
AI-assisted teaching should measure skill transfer, not answer throughput
Completion with assistance is not learning; test whether learners can retrieve and compose skills when surface form and support change.
Sep 2026 -
Coding-agent benchmarks should grade repair, not first-pass success
Measure diagnosis, repair, preserved progress, and regression control—not only whether the first patch passes.
Sep 2026 -
Grade calibration and escalation, not fluent answers
Agent education and evaluation should reward knowing when to answer, retrieve, clarify, or escalate — because fluent completion hides the metacognitive failure that matters.
Aug 2026 -
Interpretability is a teaching problem — mechanistic insight only matters if humans can act on it
Mechanistic interpretability finds circuits and features inside models, but the insight only matters if humans can act on it — making interpretability fundamentally a teaching and interface design problem.
Teaching -
Prompt training ages badly; evaluation literacy compounds
Teams that can define, sample, judge, and revise quality adapt across model cycles; prompt-pattern memorization depreciates with each release.
Aug 2026 -
Skill transfer depends on task-distribution design, not more demonstrations
Transfer comes from task families that expose reusable structure and test adaptation under variation.
Sep 2026 -
Teach evaluation through disagreement, not answer keys
Disputed judgments expose evidence choices, rubric ambiguity, uncertainty, and real value trade-offs that canonical scores conceal.
Sep 2026 -
Teaching evaluation literacy requires preserving judge disagreement
Criterion-level variance, rationales, and judge identity teach learners what consensus hides about ambiguity, bias, and measurement.
Teaching & training Sep 2026 -
Teaching skill composition requires grading the seams, not component recall
Compositional competence appears at routing, interface, validation, and recovery boundaries, so those seams must be taught and graded directly.
Teaching & training Sep 2026
No notes match that filter.