Product and performance
Product and performance
AI Evals (Evaluations)
AI evals are automated or human frameworks that measure model accuracy, bias, safety, and task performance before and after you ship.
Compute
Compute is the processing power (GPUs, TPUs, cloud instances) used to train models and run inference when users generate, classify, or embed content.
Fine-Tuning
Fine-tuning adapts a base model to your domain, tone, or task by training on curated examples, beyond what a system prompt alone can reliably enforce.
GEO (Generative Engine Optimization)
GEO (generative engine optimization) is the practice of shaping content and structure so AI answer engines (ChatGPT, Perplexity, Gemini, Claude) cite and summarize your product accurately.
Inference
Inference is running a trained model on new inputs to produce outputs: the live “prediction” step users experience as chat, classify, or generate.
Latency
Latency is the delay between a user action and a usable AI response: time to first token, time to complete answer, or time to finish an agent run.
Memory
Memory is how an AI product retains user preferences, facts, or past context across sessions, beyond the single context window.
Personalization
Personalization tailors AI behavior or content to a user or segment, using memory, history, embeddings, or fine-tuned priors.
Streaming Response
Streaming is when the AI sends its answer incrementally as tokens generate, instead of waiting for the full reply.
Token Burn Rate
Token burn rate is how fast a product consumes tokens over time: per request, per user session, or per agent run.
Vibe Coding
Vibe coding is iterative building with AI coding tools (Cursor, Claude Code, Lovable, etc.) where natural language steers rapid prototypes,“make it feel calmer,” “add citation chips.”