IntentLens
Four AI agents that read thousands of customer reviews and tell you what people actually wanted, with another LLM acting as the judge of how good the answers are.
The idea
The problem
Nobody reads 5,000 reviews. A star rating says how a customer felt, but not what they were trying to do, and a single prompt to one big model tends to blur the two.
The answer
Split the work across a small crew of agents, each with one clear role, and let a supervisor keep them honest. Then measure the result instead of trusting it.
How the crew works
A CrewAI team. Roles are narrow on purpose, and the supervisor sees everything.
The scraper agent collects the reviews.
The intent analyser reads each review and works out what the customer meant.
The search agent looks things up on the web when a review needs context.
The supervisor coordinates the others, checks their output, and assembles the final analysis.
An LLM judge scores the answers, so quality is a number.
Meet the agents
Four agents, four jobs. None of them does anyone else's.
Scraper
Goes and gets the reviews. It does not interpret them.
Intent analyser
Reads a review and says what the customer was really after.
Search agent
Goes to the web for context that the review itself does not contain.
Supervisor
Directs the crew, rejects weak work, and puts the final answer together.
Design choices
Why it is built as a crew and not as one prompt.
A narrow role is easier to prompt, easier to test, and easier to replace than one agent doing everything.
Nothing leaves the crew unchecked. The supervisor is the one place where quality is enforced.
LLM-as-judge evaluation turns 'it looks fine' into a score you can compare between runs.
A Streamlit app lets you explore the results instead of reading logs.
Built with
| Agents | CrewAI |
|---|---|
| Model access | Groq API |
| Interface | Streamlit |
| Evaluation | LLM-as-judge |