Abdul HaqueHire me
Back to my work
RSC-05 · Aug 2024 – Apr 2025 · research

GraphRAG-Causal

My final-year research. A framework that turns news headlines into causal knowledge graphs, then uses the nearest graphs to teach an LLM what causes and effects look like. Published on arXiv.

CrewAI FlowsNeo4jFastAPIarXiv
F1 score82.88%on CausalNewsCorpus
Accuracy80%on par with fine-tuned BERT-Large
Annotated sentences1,023cause, trigger, effect

The idea

The problem

News is full of causes and effects, and many are never stated outright. Classic NLP struggles with these implicit links, especially when there is very little labelled data.

The answer

Skip the fine-tuning. Turn labelled headlines into a graph, find the past events most like a new sentence, and show them to an LLM as examples. The model learns from its neighbours.

The three-stage pipeline

Annotate, retrieve, infer. The same three stages the paper describes.

ANNOTATERETRIEVEINFERACT
1. ANNOTATE

Headlines are labelled with cause, trigger, and effect, written as XML tags, and converted into causal graphs.

2. RETRIEVE

Graphs and their embeddings are stored in Neo4j. A hybrid Cypher query finds events that are close in meaning and in graph shape, using multi-hop traversal.

3. INFER

The top K retrieved graphs (5, 10, 15 or 20) go into a few-shot prompt written in XML. The LLM classifies and tags the causal relationship.

4. ACT

An agentic extension in CrewAI Flows plugs in the graph database and web search, for controllable causal inference.

What a causal graph looks like

An illustration of the idea, not a line from the dataset. Every labelled sentence becomes a small graph like this, and the graphs link up into a bigger one.

Heavy rainfloodsroad closeddelivery delays

Multi-hop traversal follows paths like this one, so a new sentence can match an old event that shares its structure and not only its words.

The numbers

Measured on CausalNewsCorpus, with the examples retrieved from the graph.

Result

82.88% F1 and 80% accuracy, on par with fine-tuned BERT-Large

Paper abstract

reports an F1 of 82.1% on causal classification with just 20 few-shot examples

Dataset

1,023 news sentences labelled with cause, trigger, and effect

Retrieved examples

Top K of 5, 10, 15, and 20

Models tried

DeepSeek distill Llama 70B and Llama 4 (Maverick)

Why it matters

News reliability, misinformation checks, and policy analysis

Why it works

Four choices that carry the result.

Few labels

Graph retrieval replaces fine-tuning, so the method works in low-data settings.

Meaning and shape

Embeddings match what a sentence says. Graph structure matches how its causes and effects connect. The hybrid query uses both.

Consistent answers

XML prompts give the model a fixed shape to follow, so the output is easy to check and parse.

More than a label

A GUI shows the causal graph, and the CrewAI Flows extension lets an agent query the graph and search the web.

Built with

FrameworkCrewAI Flows, FastAPI
GraphNeo4j with Cypher, plus embeddings
ModelsDeepSeek distill Llama 70B, Llama 4 (Maverick)
ToolingDocker, notebooks for loading data and embeddings, a GUI to explore results