The DV Intelligence Stack
The DV Intelligence Stack
Chip design verification is one of the few engineering disciplines where the gap between what AI can theoretically do and what it actually does in practice is still very wide. Not because the models aren’t capable, frontier LLMs can reason about SystemVerilog with genuine depth. The gap is architectural. Session-based agents, generic retrieval, and context-window-sized thinking don’t match the temporal scale and structural complexity of real DV work.
This series documents the engineering required to close that gap. Each article is self-contained but builds on the previous one. The through-line: a persistent, structured knowledge layer built from the codebase itself, with LLMs invoked on-demand as reasoners, not as orchestrators.
Why LLMs fail at DV
Current LLMs struggle with SV and UVM in predictable, fixable ways. Here’s what’s broken and why it matters.
The keyword ‘virtual’ means four unrelated things in SV and UVM. A precise map of each, and what happens when an LLM gets it wrong.
Teaching a model to read SystemVerilog
Fine-tuning GraphCodeBERT for SV/UVM named entity recognition, the extraction layer that makes everything else possible.
The NER model cuts 6,000 token codebases to 200-token entity records. Does that compression finally make local models usable for DV reasoning?
What fine-tuning GraphCodeBERT for SV/UVM NER involved, what worked, and the limits of general-purpose models adapted to a specialist domain.
Closing the UVM gap GraphCodeBERT can’t bridge: what specialized architectures cost, and the pre-training strategy behind cross-file inheritance resolution.
The knowledge graph and daemon architecture
Building a typed knowledge graph from extracted design entities. What relationships matter in UVM and RTL, how to represent them, and what provenance requires.
Why session-based LLM agents fail on 20-hour simulation runs, and the daemon-first architecture that works instead.
How ontology-aware graph retrieval produces causally complete subgraphs for LLM reasoning, and why typed traversal beats generic RAG for tractable bug classes.
Deploying LLM agents against a typed subgraph for iterative structural reasoning: the traversal pattern for deadlocks, error masking, and spec compliance bugs.
Putting the reasoning layer to work
On-demand reasoning for coverage analysis, assertion explanation, and regression triage, where single-call LLM reasoning outperforms structured traversal.
Using RTL change impact analysis and graph traversal to predict which tests will catch a given commit, cutting simulation time without raising risk.
What becomes possible when the design graph is live and queryable: natural language and structured queries against the verification environment.
Packaging the full stack, extractor, daemon, graph store, query interface, so the team can use it without touching training infra or Python environments.
The full picture
Last updated on