Applied Research · Grid Infrastructure

From Panel Image to Wiring Compliance

An explainable graph-based verification framework

Cassandra Rounds · Yekaterina Mijatovic · Ishita Jain · Poornima Joshi · Alexander Berry · Ethan Gueck

Overview

The problem

Utilities need to verify that field installations match approved engineering designs. Today, that can mean manually reviewing hundreds of field photos and comparing them against schematics.

The result is a slow process that consumes hundreds of expert engineering hours, and discrepancies can still go unnoticed until late.

We are exploring an explainable, automated approach that turns both field installations and engineering designs into a common representation, then compares them for compliance.

Why this is hard

01

Physical drift

01

Assets change over time. Components are repaired, replaced, or moved, so the field installation may no longer match its original design.

02

Records debt

02

Drawings get mislabeled and photographs get lost — and there are dozens of photographs per cabinet. Keeping it all organized and current is work that never ends.

03

Knowledge in people

03

The devil is in the details, and those details live in the experience and memory of a small group of experts. They are not written down in any manual, so the knowledge stays scattered.

04

Deferred cost

04

A discrepancy caught during inspection may be easy to fix. The same issue discovered later can mean rework, outages, or safety risk.

Approach

Framework at a glance

Figure 0 · the idea
Approved design · schematic BUS W1 W2 W3 GND As-built installation · field photos BUS W1 W2 W3 GND Attributed graph comparison Ground bond absent · critical Way 3 not landed · major Edit path → compliance score
Both the design and the installation become attributed graphs over one ontology. What separates them is the finding.

Diagrams scroll sideways on small screens.

We present a framework that encodes both the as-built installation and its engineering schematic as attributed graphs over a shared ontology, and verifies compliance in two layers: deterministic rules for universal installation invariants, and severity-weighted graph edit distance comparing the installation against its specific design — the distance doubles as a compliance score, the edit path as a findings list. Every finding cites the affected component, its severity, and the governing standard — auditable and training-free, not a black-box score. Schematic graphs are extracted automatically from production CAD drawings; the computer-vision front-end that produces as-built graphs from field photographs enters the framework through a fixed interface contract.

Data sharing and security practices

The data for this project was provided by an electric utility under a non-disclosure agreement (NDA), which meant the utility, not the research team, set the requirements for how it could be handled. Utility infrastructure records can contain sensitive CEII and reveal how critical grid assets are configured and where they may be vulnerable, so all development had to take place in a secure local environment. No protected data was uploaded to cloud platforms or processed by externally hosted large language models (e.g., ChatGPT, Claude); all model inference ran on locally hosted models. Any data that needed to leave the secure environment, for any reason, was first thoroughly scrubbed so that nothing could be traced back to specific assets, locations, or the utility. These requirements shaped the system architecture as much as model performance did, and they are typical of critical-infrastructure sectors governed by frameworks such as NERC CIP and NIST SP 800-53.

Two further requirements followed from the regulated context. The first is explainability: engineers and regulators must be able to understand why the system flags a discrepancy. For that reason, the machine learning and vision-language components handle only unstructured inputs, and compliance decisions are made by deterministic graph comparison with full source provenance.

Auditability was the second key requirement. Regulation of AI in the utility sector is still limited, but we expect it to expand. We argue that models intended for regulated environments should be built from the outset to withstand audit, meaning they can be broken into components that can each be inspected, explained, and independently tested. Development teams should plan for this as a baseline expectation. For practitioners building models for critical infrastructure, these constraints are not obstacles to work around but design requirements to meet from the start.

Pipeline

How it runs

Three stages are documented below: the one-time model development plan for the photo-to-graph stage; inference on a new set of site photos; and graph comparison against the design.

built and measured planned or pending upstream work stage output

Model development plan

Figure 1 · one time
DETECTOR TRACK — BUILT AND MEASUREDPhoto archivesite visits · multi-anglelocal onlyAnnotateLabel StudioDataset build15 classesTrain detectorarchitectures stillbeing comparedEvaluateMLflow → DagsHubmore labels, new classesthe trained detector supplies the crops for both tracks belowPIGTAIL CLASSIFIER TRACK — PLANNEDGenerate cropselbow crops from detectionsManual sortwire_present / no_wireTrain classifiermodel TBDEvaluateShip weightsOCR TRACK — PLANNEDGenerate text cropsmicarta_tag · op_numberother_textTranscribe sampleground truth for measurementEvaluate OCR enginesengines TBDSelect or fine-tune

Scroll the figure sideways to see the full width.

The detector track is complete and measured, but tuning and refinements are still ongoing. The classifier and OCR tracks come next. Both depend on a working detector to supply their crops, so they run after the detector track, not alongside it. The OCR track selects and measures rather than trains — but it still needs transcribed ground truth, without which there is no way to tell whether the extracted text is right. Each stage of this pipeline is also a pipeline, with different models responsible for different tasks. Some can be worked on independently; others depend on the results of the step before. Detection is the foundation of everything downstream, so it is where the effort goes first — nothing later in the pipeline can be better than the detections it is handed.

Inference on new photos

Figure 2 · every run
Site photosone cabinet · N anglesDetector15 component classesCrop + classifyelbow crops · pigtail wire presenceCrop text regionsmicarta_tag · op_number · other_textOCRequipment identifiers · added contextPer-photo findingsclass + box + wire state + textPer-cabinet findingsone row per physical partN photos collapse to one cabinetGraph buildas-built graphComparison stagesee Figure 3

Trained weights are applied to new photos — no annotator, no training. Detection and classification reason about a single photo, while a graph is about a whole cabinet, so repeated observations of the same physical part have to collapse into one cabinet-level representation.

Graph comparison and compliance

Figure 3 · every run
Design schematicproduction CAD · DXFAs-built graphfrom Figure 2Schematic graphautomatic extraction · no modelOntology alignmentshared node + edge attributesLayer 1 — deterministic rulesuniversal installation invariantsLayer 2 — severity-weighted GEDasymmetric edit costsHungarian bipartitebaseline solverSimulated annealingQUBO formulationQAOA · Qiskitbehaviour characterizationEdit path + distancethe deviations, ranked by severityFindingscomponent · severity · governing standardCompliance scoreExplainable reportauditable · training-free

The design side is deterministic: production CAD becomes a schematic graph without a model in the loop. Comparison runs in two layers. Rules catch what is wrong regardless of design; graph edit distance catches what is wrong for this design. Edit costs are asymmetric so the distance carries severity rather than merely counting differences. The three solvers are benchmarked on generated cabinet graphs; the quantum track characterizes QAOA behavior rather than claiming a speedup.

Worked example

Two traces, end to end

Two traces through the framework, end to end, on material cleared for public release.

Trace A — Photograph → annotations → JSON → graph

A publicly available cabinet photo carried through detection, structured output, and the as-built graph.

Original image
Field photograph of a dead-front switching cabinet interior showing three phases.
Labeled image
The same photograph with detection boxes over elbows, pigtails, and the ground bus.
Parsed JSON
{
  "name": "as_built_431_from_photo",
  "cabinet_attrs": {
    "manufacturer": "S&C",
    "op_number": "1234-B12",
    "cabinet_type": "dead_front",
    "bushing_type": "dual",
    "kv_level": 14.4,
    "supervisory": false,
    "bay_configuration": 4,
    "circuit_config": "circuit_isolated",
    "model_number": "431",
    "orientation": []
  },
  "nodes": {
    ...
    "BAY_4": {
      "node_type": "bay",
      "phase": "none",
      "label": "BAY_4",
      "confidence": 1.0,
      "bay_class": "fused",
      "bay_type": "TX",
      "bay_number": 4,
      "bay_position": 4,
      "num_conduits": 1,
      "num_wires": 3,
      "num_bushings": 3
    },
    ...
  },
  "edges": {
    "BAY_1|GND_BUS_1": {
      "edge_type": "contains",
      "phase": "none",
      "confidence": 1.0
    },
    ...
  }
}

Trace B — Schematic → DXF → JSON → graph

One scrubbed or generated schematic carried through CAD extraction, structured output, and the design graph.

Schematic
Box pad details top view and electrical layout of the switching cabinet.
Box pad details and electrical layout — the production drawing the design graph is read from.
DXF entities
...
  1
______ ABC
  7
ENG
 72
2
 11
9.88
 21
12.379999999999999
 31
0.0
100
AcDbText
  0
TEXT
  5
...
Parsed JSON
{
  "name": "schematic_431_8476B11",
  "cabinet_attrs": {
    "manufacturer": "S&C",
    "op_number": "1234-B12",
    "cabinet_type": "dead_front",
    "bushing_type": "dual",
    "kv_level": 14.4,
    "supervisory": false,
    "bay_configuration": 4,
    "circuit_config": "circuit_isolated",
    "model_number": "431",
    "orientation": []
  },
  "nodes": {
    ...
    "BAY_4": {
      "node_type": "bay",
      "phase": "none",
      "label": "BAY_4",
      "confidence": 1.0,
      "bay_class": "fused",
      "bay_type": "TX",
      "bay_number": 4,
      "bay_position": 4,
      "num_conduits": 1,
      "num_wires": 3,
      "num_bushings": 3
    },
    "FI_4A": {
      "node_type": "fault_indicator",
      "phase": "A",
      "label": "FI_4A",
      "confidence": 1.0
    },
    ...
  },
  "edges": {
    "BAY_1|GND_BUS_1": {
      "edge_type": "contains",
      "phase": "none",
      "confidence": 1.0
    },
    ...
    "W_4A|FI_4A": {
      "edge_type": "fault_indicator_on",
      "phase": "A",
      "confidence": 1.0
    },
    ...
  }
}
Comparison · graph to graph
Side-by-side schematic and as-built graphs with compliance metrics and highlighted differences.
431 Cabinet — DXF design vs. field photo. GED 51.50 · 12 issues · 3 high · 0 critical. The boxed region is the fused bay: the design carries a fault indicator the as-built does not.

Reference

S&C Spec Bulletin 665-31 — specification-bulletin-665-31.pdf

Arcade

Play the problem

Two small games built on the same ideas as the framework — one about wiring a cabinet, one about reading the graph. Both open in their own page.

Team

Who built it

CR

Cassandra Rounds

LinkedIn

Cassandra Rounds

Cassandra Rounds is a senior engineer with 15 years of experience in transmission and distribution engineering within the electric utility industry, spanning design, field operations, and risk analysis. Outside of standard work responsibilities, she conducts independent research applying machine learning to infrastructure verification problems, combining deep domain expertise in utility systems with applied skills in computer vision and graph-based modeling. Cassandra's work focuses on building explainable AI systems that bridge engineering documentation and field reality. She is particularly interested in the intersection of domain expertise and applied machine learning, and in developing tools that earn trust from the operational teams who must rely on them.

YM

Yekaterina Mijatovic

LinkedIn

Yekaterina Mijatovic

Yekaterina (Katya) Mijatovic is a Principal Data Scientist at Data Society Group in Washington, DC, leading technical teams with focus in scientific computing, machine learning, and AI applications. A data science generalist in the past, Katya has focused on the Energy and Utilities industry, developing custom solutions and data products with engineers and their business needs in mind. Her experience spans multiple programming languages and data science tools. Katya finds joy in cooking for her family and having guests over. She may or may not have gotten in over her head with some home improvement projects in the past. This year Katya has finally started taking ice skating lessons, it only took her 30 years to take the plunge.

IJ

Ishita Jain

LinkedIn

Ishita Jain

Ishita Jain is a machine learning engineer with over five years of experience building data science, artificial intelligence, and machine learning solutions across the energy sector and enterprise applications. Her work spans end-to-end data platforms, document intelligence, large language model (LLM)-powered applications, predictive analytics, and scalable cloud-based systems, with a focus on transforming complex data into practical solutions. She has contributed to projects that improve power grid reliability through automated asset health assessment, as well as AI-powered recommendation systems and intelligent decision-support applications. Ishita is particularly interested in artificial intelligence and agentic applications, with a focus on building scalable AI systems that enhance decision-making, automate complex workflows, and create meaningful real-world impact.

PJ

Poornima Joshi

LinkedIn

Poornima Joshi

Poornima Joshi is a Lead Data Scientist at Data Society Group, specializing in artificial intelligence, machine learning, and advanced analytics. She helps organizations transform data into actionable insights and scalable solutions, with a focus on practical AI applications that deliver measurable business value. Passionate about making AI accessible and impactful, Poornima works with teams to move from experimentation to successful real-world implementation and adoption.

AB

Alexander Berry

LinkedIn

Alexander Berry

As a senior data scientist at Data Society Group, Alex Berry bridges the gap between complex technical systems and strategic business goals. With a Master's degree in Data Science from Brown University, Alex specializes in time-series forecasting, machine learning, and cloud-native MLOps within the energy and utility sectors. He designs scalable end-to-end forecasting pipelines and intuitive analytics dashboards that help organizations optimize operations and anticipate capacity bottlenecks. When he's not building production-grade predictive models, Alex spends his winters on the ski slopes and his summers reading historical fiction or playing Mahjong at the beach.

EG

Ethan Gueck

LinkedIn

Ethan Gueck

Ethan obtained his B.S. in Network Engineering and Security and went on to complete his Master's in Data Science. His experience in recent years has been primarily in the telecom and utility industries, where he has built operational, financial, and risk based data pipelines and analysis, utilizing foundational analytics concepts as well as implementing more advanced modeling and mathematical solution sets.

Contact

Get in touch with the team

Questions about the method, the data handling, or working together on a similar verification problem — tell us a little about you and we will get back to you.

Submissions go to the team through Netlify Forms. We do not share your details with anyone.