Streaming is the default temperature of agent-era data

Data platform & strategy Growing Planted Aug 2026 · Tended Sep 2026

For data that drives immediate agent action, I start with streaming and make batch justify its delay. Not because every agent needs millisecond updates, but because the useful lifetime of a fact is set by the decision it supports. A nightly dashboard can declare its reporting period; an automated action needs an equally explicit freshness contract. The agent that offers a discount for a cart abandoned four hours ago — after the customer already bought — is acting on a world that no longer exists. Faster inference won't repair that. The question isn't whether the pipeline is called real-time. It's whether the evidence is still valid when the action happens.

The reverse-ETL literature makes the failure concrete: event to warehouse to model to sync to action, with delay accumulating at every handoff. CDP.com's “open feedback loop” critique comes from a packaged-CDP vendor, not a neutral verdict on every warehouse-native architecture. But the architectural question survives the sales pitch: does feedback return before the next decision needs it? I would separate that deadline from the model's retraining schedule. An agent may need a purchase event immediately without needing new model weights immediately. The same distinction matters on the read path. Pinot, Druid, and StarRocks target fast analytical queries over continuously arriving data; a sub-second query still tells you nothing by itself about how recently its inputs arrived. Query latency and data freshness are separate promises.

The substrate is already moving

What makes streaming a credible default is a wider choice of cost and latency trade-offs, not the disappearance of either. Object-storage-native designs such as WarpStream and AutoMQ move durable streaming data toward the same storage substrate as the lakehouse, reducing dependence on broker-local disks and cross-zone broker replication. KIP-1150 carries that direction into Kafka's design process; it is an umbrella proposal, not itself a shipped implementation. NATS makes a different boundary visible: its low-latency Core pub/sub is at-most-once and non-persistent, while JetStream adds persistence, acknowledgments, and replay. Message deduplication does not make an external tool effect exactly-once. A worker can complete the effect and crash before acknowledging it; the destination still needs idempotent handling. I want those guarantees named separately, not compressed into the word “streaming.”

The honest counter-case is that "default" is not "universal," and the streaming world itself supplies the caveat. Diskless designs make latency an explicit buffering-window trade-off — batching writes to object storage can add hundreds of milliseconds, acceptable for some telemetry and outside the budget of some fraud decisions. The workload name doesn't settle that choice; the action deadline does. Not every stream needs to be hot, and paying for heat you don't need is its own failure mode. Genuinely batch workloads remain, too: model training, monthly reconciliation, anything where recency bias runs the other way and last quarter matters more than last second. Streaming also carries real operational weight, which is exactly why event sourcing must earn its complexity rather than be adopted on vibes. The claim is about the default, not the extremes: when a new dataset arrives, the old presumption was batch-unless-proven-urgent. Agent-era consumers invert it — current-unless-proven-tolerant. Staleness becomes something you justify, budget, and declare, because the consumer can no longer smell it.