The semantic layer is what your LLM actually queries
I no longer think the production boundary for natural-language analytics is “can the model write SQL?” It is “can the model name a governed intent that the platform can validate before it touches data?” Raw-schema text-to-SQL asks the model to rediscover the company on every question: which revenue measure is authoritative, whether customers means accounts or people, which join preserves the intended grain, and how “California” is stored. Schema linking alone accounts for 30–40% of failures in the source corpus. Fluent SQL is therefore a misleading milestone. The semantic layer should be the LLM’s compiler target, not extra prose in its prompt.
I would make the model emit a typed semantic request: metric, entity, dimensions, grain, time window, filters, and requested ordering. “Margin for Audio last quarter” should resolve to a certified margin_pct object, the canonical Audio entity, and an explicit fiscal period—not to a free-form join assembled from table names. The semantic runtime can then reject an unknown metric, an illegal dimension, a non-additive roll-up, or an ambiguous entity before compilation. Only after that validation should MetricFlow, a warehouse-native metric view, Cube, or another engine bind the logical request to physical tables and generate dialect-specific SQL. This changes the model’s job from inventing a query to selecting and composing valid objects.
The operational path matters as much as the abstraction. For a large estate, I would retrieve only the relevant semantic objects and their descriptions, then resolve user literals against actual database values—“California” to CA, for example—rather than trusting lexical resemblance. Canonical entity IDs should cross the metric model and any knowledge graph so that “Audio,” a product-family code, and a catalog node do not become three identities. The generated SQL still passes an AST validator, declared-schema and join checks, and read-only enforcement; execution still needs row limits, timeouts, and CPU or memory quotas. A semantic layer narrows the valid search space. It does not turn generated SQL into a privileged workload.
I would also separate structure, meaning, and trust. Structure says what data exists and how it connects. Meaning defines metrics, business rules, synonyms, grains, and legal joins. Trust records which definition is certified, who owns it, what lineage and access policy apply, when it was last checked, and which golden queries or corrections have validated it. That third layer cannot be inferred from a metric formula. A portable semantic definition may travel between tools while its freshness, provenance, approval, and verification history do not. For an agent, that is a dangerous gap: a governed definition without current trust evidence can produce a wrong answer that looks unusually authoritative.
So I would return an answer with a decision receipt, not just a chart: semantic-object identifiers and versions, the resolved entity IDs and filters, compiled SQL, policy decision, execution limits, lineage, certification and freshness status, and any warning or truncation. I would log the natural-language question, typed request, SQL, result status, latency, errors, and user correction. Those records become the evaluation set. Execution accuracy is necessary, but I would also run query-variance tests: three to five meaning-preserving paraphrases should resolve to equivalent semantic requests and results. If “revenue,” “sales,” and “income” silently select different governed objects, the interface is not robust even when each query executes.
Knowledge graphs extend this design only where relationships carry operational meaning. Metrics and governed entities answer quantitative questions; a graph can add policy, document, event, and multi-hop relationships, with canonical IDs and provenance-scoped assertions. I would not make an ontology the default destination for every query. I would add it when path traversal or scoped reasoning changes the answer, while keeping metric computation in the semantic engine.
The precise concession is that genuinely exploratory questions over unmodeled data still require raw-schema access, clarification, and human review; a small, clean five-table schema may not justify this machinery. But repeated business questions should not pay that ambiguity tax forever. The durable interface is natural language → typed semantic intent → governed compilation → constrained execution → evidence-bearing answer. The LLM did not remove modeling. It made the missing contract executable.