Link Prediction and Outlier Detection in Industrial Knowledge Graphs
Companies like Cognite collect information about maintenance work, machines and sensors in one large network, called a knowledge graph. If a connection in that network is wrong or missing, everything built on top of it inherits the mistake. Together with Cognite, we are using machine learning to find those errors automatically.
- Machine Learning Engineer
- Student team, in collaboration with Cognite.
- Sep 2026 – present
- Ongoing
- Python · rdflib · Node2Vec · scikit-learn · UMAP · matplotlib
Problem
Cognite builds industrial knowledge graphs: networks where every maintenance operation, piece of equipment, maintenance order, sensor time series and asset is a node, and the relationships between them are edges. An operation should point to the equipment it is performed on and to the order it belongs to; equipment belongs to assets in a physical hierarchy.

When these links are wrong or missing, every analysis and AI system built on top of the graph inherits the error. Rule-based validation catches the problems someone thought of in advance, but large graphs that are merged from many source systems contain errors nobody anticipated.
The question is whether we can learn what a normal part of the graph looks like, and then use that to answer two things:
- Outlier detection: which entities look structurally unusual compared with the rest?
- Link prediction: which relationships are probably missing, and what should they point to?
There are no labels saying which links are wrong, so the methods have to work without a ground truth, and results must be explainable to the people who own the data.
Technical skills
- Knowledge graphs and RDF data (rdflib)
- Exploratory data analysis and data-quality assessment
- Graph analysis: degree distributions, connected components, relationship profiles
- Statistical outlier detection (rarity scores, IQR, robust z-scores)
- Link prediction and evaluation with ranking metrics (Hits@k)
- Graph embeddings with Node2Vec, and dimensionality reduction (PCA, t-SNE, UMAP)
- Reproducible Python pipelines with command-line scripts
My contribution
- Analysed the structure and quality of the graph, and documented what is normal, what is missing and which questions only the data owner can answer.
- Built the pipeline from raw RDF data to a graph and to node embeddings, with scripts the team can reuse.
- Developed statistical baselines for outlier detection and link prediction, so later machine learning models have something concrete to beat.
- Helped move the team from data exploration towards outlier detection and link prediction.


Approach
Understand the data first
Before modelling, map how the node types connect, what is missing and what looks inconsistent.
Represent the graph
Turn RDF triples into a graph that algorithms can work with, and remove generic hub nodes that would dominate everything.
Set simple baselines
Statistical outlier scores and simple link-prediction heuristics, with clear criteria for what a better method must achieve.
Learn from structure
Node2Vec embeddings turn each node into a vector based on its neighbourhood, so similar nodes end up close together.
Next: machine learning models
Embedding-based outlier detection and link prediction, compared against the baselines. This is the phase we are in now.