Skip to content
(02/03)  Case studyAI · Production RAG agent

An AI that talks like a real creator.

A production RAG agent with a versioned personality, a knowledge base from his whole video library, and live alerts that can overrule the model.

Role
Designed & built end to end
Year
2026
Context
Associated with ClanFlare
Links
Private · NDA

Next.js / n8n / RAG / Pinecone / Cohere / Upstash / Redis

private client buildAnonymized
Enterprise AI Persona Chatbot: recording of the app
95%+
Persona adherence
~20%
Lower latency
2,000+
Users
50+ hrs
Media indexed

The goal

Not a FAQ bot with his face on it.

His audience should talk to it like it's him, and not be able to tell the difference.

The shape of it

One product, two minds.

01

A thin interface

A Next.js chat UI: render the answer, embed the source clip, stay out of the way.

02

A separate brain

An n8n agent that retrieves, wears the persona, calls the model, and enforces alerts. The whole agent lives behind one webhook.

03

Purpose-built memory

Two vector stores + Redis, each chosen for a specific access pattern, not by reflex.

private client buildPrivate
Chat interface answering in the creator's voice, with the source clip and live alerts
Fig. 01Answers in his voice (short, blunt, Hinglish) with the exact clip it's drawing from, and the desk's live alerts alongside.

Live alerts

Where a human sets the truth.

An admin posts a market alert; the agent treats it as reality and overrules its own take. Stale-but-confident is the enemy: a fresh, human-posted alert beats the model's own opinion.

private client buildPrivate
An answer deferring to a live desk alert on Gold
Fig. 02A human-posted alert on Gold is live, so the persona defers to it. The alert outranks the model's earlier read.
private client buildPrivate
The alert console: post, schedule and expire desk alerts
Fig. 03The alert console: post, schedule and expire. Every write is checked against an admin secret on the server.
System architecture: chat and admin UIs → Next.js API → n8n agent → OpenRouter, with Upstash Vector, Pinecone and Redis
Diagram· scroll →Interface separated from intelligence: Next.js → n8n agent → OpenRouter, with a two-tier retrieval layer.

Persona engineering

A system prompt, eleven times over.

01

Brevity over essays

2–4 sentences, no lists. Real traders are blunt.

02

A noise filter

Off-topic questions get an in-character brush-off, not a tutorial.

03

Tone modulation

Patient teacher for learners; firm only with shortcut-seekers, tuned to 95% adherence.

Decisions, not defaults

Why it's built this way.

01

Two-tier retrieval

Upstash Vector for the knowledge base; Pinecone + Cohere for time-sensitive alerts. HyDE query rewriting cut latency ~20%.

02

Alert override

A fresh, human-posted alert beats the model's own opinion, with an expiry so it can't go stale.

03

Redis sessions

Per-user state, isolated: 2,000+ conversations that never cross wires, with sub-millisecond reads.

Semantic dilution: keyword search buries the right clip at rank 6; searching with the full question plus HyDE puts it first
Diagram· scroll →Semantic dilution: repeated themes buried the one clip that answered the question. Searching with the full question plus a hypothetical answer (HyDE) put it on top.
One Redis key per user keeps each conversation's context isolated
Diagram· scroll →One Redis key per user: a stateless app and agent pass the sessionId, so conversations stay isolated and reads stay fast.

Mobile

The same desk, sized for a phone.

The chat on a phone: a scripted demo conversation
The conversation
A reply with its source clip on a phone
Source clip in the reply
Live alerts as a bottom sheet on a phone
Live alerts sheet
StackWhy each piece

Next.js 16 · React 19

Server components keep secrets and data access on the server.

n8n · OpenRouter

Visual orchestration; model-agnostic LLM access you can swap without a deploy.

Upstash Vector · Pinecone · Cohere

Two-tier RAG, each path tuned for its job.

Upstash Redis

Hot, ephemeral, per-user state.

Your turn

Want onelike this?

Tell me what you're building. You'll hear back within 24 hours, with a fixed quote before any work starts.