Table of Contents
- Building Catalyst: Architecture of a Schema-First Conversational Spreadsheet Intelligence Agent
- 1. Traditional Chat Threads vs. The Catalyst Protocol
- 2. The Engine: Hybrid Agentic Architecture
- 3. Client-Side Execution Sandbox & The Self-Healing Compiler
- 4. Multi-Modal Spreadsheet Intelligence in Action
- 5. The Headless FastAPI ReAct+ Engine (/api)
- 6. Enterprise LLMOps, Telemetry & Observability
- 7. Technology Stack & Edge Architecture
- 8. Key Takeaways from Building an Autonomous Data Agent
- Try Catalyst & Explore the Code
Building Catalyst: Architecture of a Schema-First Conversational Spreadsheet Intelligence Agent
Spreadsheets are the bedrock of modern global enterprise operations. From financial budgets and supply chain matrices to sales performance and customer telemetry, millions of mission-critical decisions are calculated across tabular grids every single day.
When Large Language Models emerged, the natural impulse across the software industry was to build conversational chat wrappers directly over CSV files. But anyone who has ever asked raw ChatGPT, Claude, or a standard LLM thread to calculate statistical aggregates, evaluate compound month-over-month growth, or transform 10,000-row sheets quickly runs into three fatal limitations:
- Mathematical Hallucinations: LLMs do not calculate math; they predict tokens probabilistically. Asking a language model to sum 5,000 transaction rows will consistently yield plausible-looking but completely fabricated numbers.
- Context Window & Latency Choke: Passing raw 50,000-row enterprise datasets into an LLM context window blows past token limits, runs up massive API bills, and introduces 30-second response delays.
- Data Privacy & Governance Leaks: Uploading raw enterprise balance sheets or proprietary customer records to public LLM endpoints violates enterprise compliance and privacy frameworks.
To solve these foundational problems, I designed and built Catalyst (GitHub Repository), an autonomous conversational spreadsheet intelligence platform.

Global Benchmark
Ranked #10 Globally on official SpreadsheetBench V1 - Full (912) (48.57% Verified Accuracy).
Read Benchmark Deep-Dive →
In this article, I will unpack how Catalyst works under the hood: the Schema-First Code Generation protocol, the client-side execution sandbox, real-time AG Grid memory sync, the headless FastAPI ReAct+ agent engine, 8 interactive chart visualizers, Byte-Safe Volume Protection, enterprise LLMOps with Langfuse observability, and live web-to-sheet augmentation via Firecrawl.
1. Traditional Chat Threads vs. The Catalyst Protocol
Before diving into code, let’s contrast how standard AI chat wrappers handle tabular data versus Catalyst’s deterministic protocol:
| Dimension | Traditional LLM Chat Wrapper | The Catalyst Protocol |
|---|---|---|
| Mathematical Precision | High Hallucination. Approximates sums, variances, and counts probabilistically. | 100% Deterministic. Synthesizes exact JavaScript transformation logic executed in a sandboxed client runtime. |
| Data Privacy | Zero Isolation. Transmits full raw CSV and Excel row contents across public cloud endpoints. | Schema-First Privacy. Only column metadata and a tiny 3-row anonymous sample are sent to the AI. Raw records remain in private Convex storage. |
| Dataset Size Limits | Context Window Limits. Crashes or truncates rows on datasets exceeding 2,000 to 5,000 lines. | Sub-Second Performance. Smoothly processes 20,000+ rows instantly using browser-native AG Grid memory and Byte-Safe protection. |
| Web Augmentation | Static Cutoff. Limited strictly to model training weights. | Live Web Scraping. Automated agents query live search engines (Firecrawl and LangSearch) and merge real-time variables directly into grid columns. |
2. The Engine: Hybrid Agentic Architecture
The core philosophy behind Catalyst is simple: Let the LLM write code, and let a deterministic runtime execute math.
Step 1: Schema Optimization (Token & Privacy Shield)
Instead of passing the entire sheet, Catalyst computes a lightweight schema descriptor:
export interface SheetMetadata {
sheetName: string;
totalRows: number;
columns: Array<{
field: string;
type: "string" | "number" | "date" | "boolean";
}>;
sampleRows: Record<string, any>[]; // Strictly 3 representative rows
}
export function extractSchema(data: Record<string, any>[]): SheetMetadata {
if (!data || data.length === 0) return { sheetName: "Sheet1", totalRows: 0, columns: [], sampleRows: [] };
const sample = data.slice(0, 3);
const columns = Object.keys(data[0] || {}).map((key) => {
const val = data.find((r) => r[key] !== null && r[key] !== undefined)?.[key];
return {
field: key,
type: typeof val === "number" ? "number" : typeof val === "boolean" ? "boolean" : "string"
};
});
return {
sheetName: "ActiveSheet",
totalRows: data.length,
columns,
sampleRows: sample
};
}This single architectural pattern reduces LLM token consumption by over 99%, ensures that sensitive row-level enterprise records never leave the browser, and keeps query response latency under 1.2 seconds.
3. Client-Side Execution Sandbox & The Self-Healing Compiler
When the user asks:
“Calculate total revenue by region, filter out inactive accounts, and show me the top 3 performing territories.”
The LLM acts as an experienced software engineer, outputting structured execution scripts. The browser receives this script and runs it in an isolated execution sandbox.
Self-Healing Logic
Real-world datasets have typos, irregular casing (Revenue vs revenue), and null values. If the generated script encounters an uncaught reference or undefined column accessor (for example, the AI writes avgOrder instead of avgOrderValue), the Catalyst runtime catches the exception, dynamically declares the missing variable on the fly, and re-executes cleanly without crashing:
export async function executeTransformationSandbox(
code: string,
gridData: Record<string, any>[]
): Promise<{ success: boolean; data?: Record<string, any>[]; error?: string; visualResult?: any }> {
try {
// Clone dataset to guarantee zero side-effects before preview approval
const workingCopy = structuredClone(gridData);
// Sandbox wrapper providing safe mathematical utilities
const sandboxFunction = new Function(
"rows",
"Math",
`"use strict";
try {
${code}
return { success: true, transformedData: rows };
} catch (err) {
return { success: false, error: err.message };
}`
);
const result = sandboxFunction(workingCopy, Math);
return result;
} catch (error: any) {
// Variable healing wrapper: catch ReferenceErrors (e.g. typos like avgOrder vs avgOrderValue)
if (error.message && error.message.includes("is not defined")) {
const missingVar = error.message.split(" ")[0];
const healedCode = `let ${missingVar} = 0;\n` + code;
return executeTransformationSandbox(healedCode, gridData);
}
return { success: false, error: error.message };
}
}4. Multi-Modal Spreadsheet Intelligence in Action
Catalyst goes far beyond simple math queries. It is a full multi-modal spreadsheet operating environment:
Amber Cell Preview & Non-Destructive Mutations
When a transformation is computed (for example: “Capitalize all names in Column B, and for any row where Price is missing, set it to the average”), Catalyst does not blindly overwrite user data. It enters Amber Preview Mode:
- Cells undergoing changes glow in warm amber.
- Users can inspect individual diffs across the grid.
- One-click Apply or Rollback ensures complete data safety.
8 Interactive Visualization Layouts
Rather than static images, Catalyst compiles 8 distinct interactive chart types directly inside the chat feed or in executive dashboards:
- Bar & Multi-Bar Charts (Comparative totals)
- Line & Spline Visualizers (Time-series analysis)
- Area Charts (Cumulative growth curves)
- Pie & Doughnut Breakdowns (Distribution slices)
- Scatter Plots (Correlation studies)
- Radar Charts (Multi-metric performance)
- Composed & Horizontal Bar Charts (Target tracking & cross-dimensional ranking)
- Summary Metric KPI Cards (Executive rollups with dynamic tick-interval spacing)
Each chart supports full client-side tooltips, dynamic filtering, responsive resize observers, and one-click PNG/SVG export.
Autonomous Web Crawling Augmentation (Firecrawl & LangSearch)
Need to enrich customer records with their company HQ, pull live stock metrics, or lookup exchange rates? Catalyst dispatches web-crawling agents powered by Firecrawl and LangSearch, scrapes official sites, extracts structured JSON, and merges the new columns directly into your grid.
Dynamic Sheet Compilation, Cell Highlighting & State-Snapshots
- Dynamic Sheet Compilation: Conversationally command Catalyst to “Create a new sheet called Q4 Targets with mock target revenues”, and watch it compile the data array, append the tab to the workbook, and shift grid focus instantly.
- Visual Conditioning: Ask Catalyst to “Highlight all rows where COUNTRY is USA in yellow and STATUS is Shipped in light green”, applying live styling rules directly to the grid canvas.
- State-Snapshot Rollback Engine: Catalyst captures a complete matrix snapshot before any mutation, giving users instant, 0-fragmentation multi-turn Undo and Redo capability.
- Byte-Safe Volume Protection: Analyzes dataset byte sizes and safely scales operations for massive 10,000+ row sheets without browser lag or memory bloat.
arXiv Research Agent Tab
For research teams and data scientists, Catalyst includes a direct integration with the arXiv API. You can query scholarly literature (e.g., “Find the latest papers on Large Language Model reasoning architectures”), and Catalyst automatically parses titles, full abstracts, author lists, and direct PDF download links into structured spreadsheet tabs.
5. The Headless FastAPI ReAct+ Engine (/api)
In addition to the Next.js client sandbox, Catalyst houses a dedicated FastAPI Headless Agent Engine (in api/index.py). This backend service encapsulates Catalyst’s state-of-the-art ReAct+ (Reason + Act + Observe + Reflect) loop for serverless spreadsheet automation, external integrations, and document intelligence pipelines.
ReAct+ Multi-Turn Feedback Loop:
User Prompt / File
│
▼
┌──────────────┐ Generates Python Code
│ Gemini 3.1 │ ───────────────────────────────┐
│ Flash Lite │ ◄──────────────────────────┐ │
└──────────────┘ Error / Traceback Feedback │ ▼
▲ (IndexError / KeyError) ┌──────────────┐
│ │ Safe Python │
│ │ Sandbox │
│ │ (openpyxl & │
│ │ pandas) │
│ └──────┬───────┘
│ │
└─────── Auto-Reflect & Self-Correct ◄──────┘
│ (On Success)
▼
Clean Output Workbook (.xlsx / .csv)Key API Capabilities:
- Deterministic Python Transformations: Executes raw calculations in Python rather than fragile Excel formulas, ensuring 100% computational stability.
- Multi-Turn Self-Healing Feedback Loop: When script executions fail with
IndexError,KeyError, or type mismatches, the loop feeds tracebacks back to Gemini, guiding it to auto-correct within a 4-turn reflection cycle. - Multi-Key Pool Rotation: Automatically balances and rotates across multiple
GEMINI_API_KEYSto avoid 429 quota exhaustion during high-throughput batch runs. - Production Endpoints Suite (
/v1/*):POST /v1/spreadsheets/transform: Autonomous ReAct+ transformation engine on real workbooks (.xlsx,.csv).POST /v1/spreadsheets/generate: Generates complete financial models and cash-flow schedules from natural language.POST /v1/images/extract: Ingests receipts, invoices, and table photos directly into clean tabular grids.POST /v1/documents/extract: Multi-page PDF document parser generating multi-tab Excel workbooks.POST /v1/spreadsheets/audit: Scans workbooks for broken#REF!,#DIV/0!, and formula drift anomalies.POST /v1/spreadsheets/reconcile: Two-way ledger reconciliation matching bank transactions against payment processor records.
6. Enterprise LLMOps, Telemetry & Observability
Enterprise deployment requires continuous transparency into token consumption, latency, and agent reasoning traces. Catalyst integrates full LLMOps instrumentation:
1. Dual-Stream Tracing Architecture
Every call to the AI Orchestrator streams telemetry across two dedicated pipelines:
- Langfuse Cloud: Distributed trace spans capturing prompt token accounting, completion tokens, latency percentiles, and cost graphs.
- Native Convex Trace Engine (
llm_traces): Reactive in-database telemetry store powering the interactive admin observability portal.
2. Admin Observability Dashboard
- Real-Time KPIs: Total requests, token burn, average generation latency (ms), and error/quota failure rates.
- Trace Waterfall: Step-by-step trace inspection with copyable trace IDs, model parameters, and raw JSON telemetry payloads.
- Dynamic Credit Accounting Matrix: Operations are charged transparently based on actual compute and token overhead (e.g., Free for Undo/Redo, 1 credit for Conversational Q&A, 2 credits for Deep Sheet Analysis & Charts, up to 5 credits for Executive BI Dashboards).
7. Technology Stack & Edge Architecture
Catalyst was built using modern, edge-native and serverless technologies:
- Frontend: Next.js 15 (App Router, Turbopack), React 19, AG Grid Enterprise, Recharts, Framer Motion, Tailwind CSS
- Autonomous Agent API: FastAPI & Python 3.11+ (ReAct+ Multi-Turn Code Sandbox with
openpyxl&pandas) - Backend & State: Convex real-time document database & serverless mutation engine
- Authentication: Stack Auth (Cloud-native identity management)
- AI Models: Google Gemini 3.1 Flash Lite & 2.5 Pro (with multi-key pool rotation)
- LLMOps & Observability: Langfuse Cloud (OpenTelemetry Traces, Latency & Cost Tracking) + Native Convex Trace Engine
- Agent Crawling & Research: Firecrawl, LangSearch, and arXiv API
- Deployment: Vercel Edge Runtime & Serverless Python
8. Key Takeaways from Building an Autonomous Data Agent
- Never let an LLM do raw math: Always decouple reasoning (the LLM generates TypeScript/JavaScript) from execution (a deterministic JavaScript sandbox calculates the result).
- State snapshots are mandatory: Users will not trust an AI agent with their data unless they know they can roll back any edit with 100% confidence.
- Keep datasets localized: Moving large CSV files across API networks introduces unacceptable latency. Move the lightweight code to where the data already lives (in the client browser).
Try Catalyst & Explore the Code
Catalyst is live and fully open source:
- 🚀 Live Demo: catalyst.samuelolubukun.com
- 💻 GitHub Repository: samolubukun/Catalyst-The-Conversational-Spreadsheet-Agent
