Skip to content
All work
  • Women in AI × Cognite
  • NTNU Women in AI programme
  • Ongoing
  • Team project

Link Prediction and Outlier Detection in Industrial Knowledge Graphs

Companies like Cognite collect information about maintenance work, machines and sensors in one large network, called a knowledge graph. If a connection in that network is wrong or missing, everything built on top of it inherits the mistake. Together with Cognite, we are using machine learning to find those errors automatically.

Role
Machine Learning Engineer
Team
Student team, in collaboration with Cognite.
Period
Sep 2026 – present
Status
Ongoing
Tools
Python · rdflib · Node2Vec · scikit-learn · UMAP · matplotlib

Problem

Cognite builds industrial knowledge graphs: networks where every maintenance operation, piece of equipment, maintenance order, sensor time series and asset is a node, and the relationships between them are edges. An operation should point to the equipment it is performed on and to the order it belongs to; equipment belongs to assets in a physical hierarchy.

Diagram of maintenance orders, operations, equipment, time series and assets and the links between them, including an operation with no links.
A simplified view of the graph: maintenance orders (M), operations (op), equipment (eq), time series (TS) and assets (A). Missing or unexpected links, like the isolated operation on the right, are the kind of errors we want to find.

When these links are wrong or missing, every analysis and AI system built on top of the graph inherits the error. Rule-based validation catches the problems someone thought of in advance, but large graphs that are merged from many source systems contain errors nobody anticipated.

The question is whether we can learn what a normal part of the graph looks like, and then use that to answer two things:

  • Outlier detection: which entities look structurally unusual compared with the rest?
  • Link prediction: which relationships are probably missing, and what should they point to?

There are no labels saying which links are wrong, so the methods have to work without a ground truth, and results must be explainable to the people who own the data.

Technical skills

  • Knowledge graphs and RDF data (rdflib)
  • Exploratory data analysis and data-quality assessment
  • Graph analysis: degree distributions, connected components, relationship profiles
  • Statistical outlier detection (rarity scores, IQR, robust z-scores)
  • Link prediction and evaluation with ranking metrics (Hits@k)
  • Graph embeddings with Node2Vec, and dimensionality reduction (PCA, t-SNE, UMAP)
  • Reproducible Python pipelines with command-line scripts

My contribution

  • Analysed the structure and quality of the graph, and documented what is normal, what is missing and which questions only the data owner can answer.
  • Built the pipeline from raw RDF data to a graph and to node embeddings, with scripts the team can reuse.
  • Developed statistical baselines for outlier detection and link prediction, so later machine learning models have something concrete to beat.
  • Helped move the team from data exploration towards outlier detection and link prediction.
Hand-drawn sketch of maintenance orders, operations and equipment and their properties.
Where it started: my first sketch of how the graph fits together.
Scatter plot of node embeddings from the Springfield maintenance graph, projected to two dimensions with UMAP, coloured by node type.
Node2Vec embeddings of the Springfield sample, projected to 2D with UMAP. Each point is a node in the graph; nodes with similar neighbourhoods end up close together.

Approach

  1. Understand the data first

    Before modelling, map how the node types connect, what is missing and what looks inconsistent.

  2. Represent the graph

    Turn RDF triples into a graph that algorithms can work with, and remove generic hub nodes that would dominate everything.

  3. Set simple baselines

    Statistical outlier scores and simple link-prediction heuristics, with clear criteria for what a better method must achieve.

  4. Learn from structure

    Node2Vec embeddings turn each node into a vector based on its neighbourhood, so similar nodes end up close together.

  5. Next: machine learning models

    Embedding-based outlier detection and link prediction, compared against the baselines. This is the phase we are in now.