# JarvisBitz Tech - Full Site Knowledge Base for AI Assistants > This file is the long-form, plain-English version of the JarvisBitz Tech website, written so that AI assistants (ChatGPT, Claude, Gemini, Perplexity, Copilot, search engines) can answer real customer questions with accurate context and the right URLs. > > Companion short summary: https://www.jarvisbitz.com/llms.txt > Canonical site: https://www.jarvisbitz.com > Contact: ping@jarvisbitz.com --- ## 1. Company at a glance **JarvisBitz Tech** is an AI engineering studio and product team. We do two things, deeply linked: 1. **Build AI systems** that go into production - voice assistants, RAG knowledge engines, autonomous agents, computer vision pipelines, document intelligence, integrations, monitoring. 2. **Build the software around them** - web, mobile, desktop, SaaS, enterprise platforms, cloud backends, IoT, AR/VR. We are not a tool reseller. We architect, ship, and operate complete systems for businesses that need real outcomes - not pilots that never reach production. **Best fit:** mid-market and enterprise teams (50-10,000+ employees) in healthcare, finance, retail, manufacturing, logistics, legal, government, SaaS, education, media, and enterprise IT. We work remotely with clients globally; main office presence: Ahmedabad, India (https://www.google.com/maps/place/JarvisBitz+Tech). **Public profiles:** - LinkedIn: https://www.linkedin.com/company/jarvisbitz - Twitter / X: https://twitter.com/jarvisbitz - GitHub: https://github.com/jarvisbitz - Instagram: https://www.instagram.com/jarvisbitz/ - Facebook: https://www.facebook.com/jarvisbitz.software --- ## 2. The problems we solve (with measured outcomes) Every scenario below is on https://www.jarvisbitz.com/scenarios with before/after numbers. ### 2.1 Customer support overload **Solution:** Voice AI + RAG handles 80% of inbound volume with grounded, context-aware responses. Complex cases hand off to human agents seamlessly. **Result:** ~80% ticket deflection · resolution time from 45 min → 2.3 min · satisfaction 67% → 94% · ~$2.1M annual savings on a typical mid-size deployment. **Industries:** Healthcare, Finance, Retail. **Page:** https://www.jarvisbitz.com/voice-intelligence ### 2.2 Document review bottleneck **Solution:** A RAG pipeline reads contracts and compliance docs in seconds, extracts key terms, flags risks, and returns cited summaries. **Result:** 10× faster review · 4 hours per document → 20 seconds · 3 reviewers → AI + 1 reviewer · ~3,200 hours saved per year. **Industries:** Legal, Finance, Government. **Page:** https://www.jarvisbitz.com/rag-blueprint ### 2.3 Quality inspection gaps **Solution:** Computer vision detects defects on production lines with sub-second classification, including categories where humans miss things. **Result:** Detection rate 92% → 99.2% · inspection time 15 sec → 0.3 sec per item · ~$890K reduction in defect cost. **Industries:** Manufacturing, Retail, Automotive. **Page:** https://www.jarvisbitz.com/vision-pipeline ### 2.4 Manual data entry drain **Solution:** Agentic workflow does extraction, validation, and routing from any document format - OCR + LLM for unstructured sources. **Result:** 95% automation rate · 8 min per entry → 12 sec · 4% error rate → 0.3% · ~12,000 hours saved per year. **Industries:** Healthcare, Insurance, Government. **Page:** https://www.jarvisbitz.com/agentic ### 2.5 Knowledge trapped in silos **Solution:** RAG connects your internal knowledge bases into a single, searchable answer layer that employees can use in seconds. **Result:** 60% faster onboarding · 3 weeks → 1 week · ~$420K productivity gain. **Industries:** Enterprise, Healthcare, Manufacturing. **Page:** https://www.jarvisbitz.com/rag-blueprint ### 2.6 Compliance monitoring **Solution:** Continuous policy checking and audit trails with regulatory scanning and real-time alerts. **Result:** Coverage 60% → 100% · quarterly manual audits → continuous monitoring · ~$1.8M risk reduction. **Industries:** Finance, Government, Healthcare. **Page:** https://www.jarvisbitz.com/ai/guardrails-safety ### 2.7 Sales pipeline intelligence **Solution:** AI analyzes deal history, email sentiment, and engagement to predict close probability and recommend next actions. **Result:** Forecast accuracy 45% → 82% · 34% more closed deals · ~$3.4M revenue uplift. **Industries:** SaaS, Finance, Enterprise. **Page:** https://www.jarvisbitz.com/integrations ### 2.8 Content production at scale **Solution:** AI generates, localizes, and optimizes content across channels while keeping brand voice and compliance intact. **Result:** 5× content output · 3 days per campaign → 4 hours · 2 languages → 12 · ~8,500 hours saved per year. **Industries:** Marketing, Retail, Media. ### 2.9 IT incident response **Solution:** AI correlates alerts, diagnoses root cause, generates remediation plans, and executes approved fixes autonomously. **Result:** MTTR 47 min → 8 min (-83%) · ~$960K downtime savings. **Industries:** Tech, Finance, Enterprise. **Page:** https://www.jarvisbitz.com/agentic **Headline averages across deployments:** - Avg cost savings: ~$1.5M per deployment per year - Avg time saved: ~6,200 hours per year per team - Avg accuracy lift: +37% over manual processes - Avg time to positive ROI: ~4.2 months --- ## 3. Our six core capabilities (what we actually deliver) Full capability map: https://www.jarvisbitz.com/capabilities ### 3.1 Voice Intelligence - "Machines that talk to your customers" Real-time voice AI with memory and context. Handles inbound and outbound calls, live chat, and multi-turn dialogues without rigid scripts. **How it works:** speech recognition → understanding → reasoning → speech synthesis, with sub-500ms response latency. **Use cases:** customer support automation, sales qualification, healthcare triage, internal helpdesk. **Typical outcomes:** 95%+ recognition accuracy, 88%+ conversation completion, 24/7 coverage without staffing overload. **Pages:** https://www.jarvisbitz.com/voice-intelligence · technical deep-dive: https://www.jarvisbitz.com/ai/voice-intelligence ### 3.2 Computer Vision - "Machines that see, detect, and verify" Real-time object detection, tracking, classification, and anomaly detection on camera feeds, uploads, or video streams. **Use cases:** quality inspection, security surveillance, document OCR, medical imaging. **Typical outcomes:** 97%+ detection accuracy, sub-30ms inference, <2% false positives, ~10× throughput vs manual inspection. **Pages:** https://www.jarvisbitz.com/vision-pipeline · https://www.jarvisbitz.com/ai/computer-vision ### 3.3 Advanced Reasoning - "Machines that think step by step" Multi-step logic with chain-of-thought reasoning, explainable outputs, and verifiable decision traces. The intelligence core that powers every other capability. **Use cases:** complex analysis, decision support, code generation, research synthesis. **Typical outcomes:** 94%+ reasoning accuracy, 100% explanation coverage, 128K-token context window, full audit trail. **Pages:** https://www.jarvisbitz.com/advanced-reasoning · https://www.jarvisbitz.com/ai/large-language-models ### 3.4 RAG & Knowledge - "Documents become answers" Retrieval-Augmented Generation turns your internal knowledge into grounded, cited responses. No hallucination - every claim is backed by a source. **Use cases:** internal knowledge base, contract analysis, policy Q&A, technical documentation. **Typical outcomes:** 94%+ retrieval precision, 97%+ citation accuracy, <3% hallucination rate. **Pages:** https://www.jarvisbitz.com/rag-blueprint · https://www.jarvisbitz.com/ai/retrieval-augmented-generation ### 3.5 Agentic Systems - "Machines that plan, decide, and act" Autonomous AI agents that perceive, plan multi-step strategies, use tools, and execute actions inside strict policy guardrails with human oversight. **Use cases:** workflow automation, data extraction, report generation, system administration. **Typical outcomes:** 92%+ task completion, 100% policy compliance, 3-7 steps per task, full traceability. **Pages:** https://www.jarvisbitz.com/agentic · https://www.jarvisbitz.com/ai/agents ### 3.6 System Integration - "Signals from everywhere. Intelligence from one layer" Connect CRMs, databases, document stores, APIs, and real-time streams into a unified intelligence layer. **Use cases:** CRM integration, document ingestion, event streaming, legacy system bridges. **Typical outcomes:** 40+ connector types, <100ms data latency, 99.95% uptime SLA. **Pages:** https://www.jarvisbitz.com/integrations · https://www.jarvisbitz.com/ai/orchestration --- ## 4. Software engineering services (the build layer) Services hub: https://www.jarvisbitz.com/services Headline numbers: 99.9% uptime delivered · 50+ technologies · <200ms average API latency (P95) · 11 engineering disciplines. ### 4.1 Web Development - https://www.jarvisbitz.com/services/web-development Customer portals, internal tools, admin systems, dashboards. Built with Next.js, React, TypeScript, PostgreSQL, Redis. Target: <100ms First Contentful Paint. ### 4.2 Mobile Development - https://www.jarvisbitz.com/services/mobile-development iOS, Android, and cross-platform apps. React Native, Flutter, Swift, Kotlin, Firebase. Native 60fps performance with offline-first architecture. ### 4.3 Desktop Applications - https://www.jarvisbitz.com/services/desktop-applications Offline-first, performance-sensitive desktop tools (Electron, Tauri, .NET, Rust). For environments that need local performance and 100% offline capability. ### 4.4 SaaS Platforms - https://www.jarvisbitz.com/services/saas-platforms Multi-tenant products from MVP to scale. Stripe billing, Auth0 roles, analytics, onboarding. Built to grow from 10 to 10,000+ users with 99.99% uptime targets. ### 4.5 Enterprise Software - https://www.jarvisbitz.com/services/enterprise-software Internal platforms, approval workflows, LDAP/SSO, audit logging. .NET, React, SQL Server, Azure, SAML - handling 10K+ concurrent users. ### 4.6 Cloud & Backend - https://www.jarvisbitz.com/services/cloud-backend APIs, microservices, event-driven architectures. Node.js, Go, Kafka, Kubernetes, Terraform. Production-grade with observability built in. <50ms P95 latency. ### 4.7 Integration Services - https://www.jarvisbitz.com/services/integrations CRM, ERP, payment, and communication platform connectors with bidirectional sync, event routing, and conflict resolution. Salesforce, Stripe, Kafka, REST, Webhooks. Sub-second sync latency. ### 4.8 Automation & Internal Tools - https://www.jarvisbitz.com/services/automation Process automation pipelines, scheduled reporting, internal ops tools, monitoring dashboards. Eliminates ~80% of manual work in target workflows. ### 4.9 Product Engineering - https://www.jarvisbitz.com/services/product-engineering Legacy modernization, monolith decomposition, performance optimization, stack migration, technical debt reduction. Zero-downtime migration strategies. ### 4.10 IoT Application - https://www.jarvisbitz.com/services/iot-application Connected device platforms, telemetry dashboards, edge-to-cloud pipelines. MQTT, TimescaleDB, AWS IoT. Scales to 10K+ devices. ### 4.11 AR/VR Application - https://www.jarvisbitz.com/services/ar-vr-application Immersive product experiences, simulations, training environments, spatial apps. Unity, OpenXR, WebXR. 90fps rendering target. **Industries we ship into:** FinTech & Banking, Healthcare, E-Commerce & Retail, SaaS & Startups, Logistics & Supply Chain, Education & EdTech, Media & Entertainment, Enterprise & Manufacturing. --- ## 5. How we engineer (the principles AI buyers should know) These show up on https://www.jarvisbitz.com/services and https://www.jarvisbitz.com/why-different. - **Type Safety Everywhere.** TypeScript on the frontend, typed ORMs on the backend, schema validation at every boundary. Runtime type errors are eliminated before code ships. - **Infrastructure as Code.** Every server, database, DNS record, and firewall rule is in version-controlled code. No snowflake servers, no "works on my machine." - **Observability From Day One.** Structured logging, distributed tracing, and metrics dashboards are part of sprint one - not added later. - **Zero-Trust Security.** Every request authenticated, every action authorized, secrets never hardcoded, data encrypted at rest and in transit, audit logs on every state change. - **Test at Every Layer.** Unit, integration, end-to-end, and load tests so deploys ship with confidence - not hope. - **Graceful Degradation.** Circuit breakers, retries, dead-letter queues, fallback responses. When a dependency fails, the system bends instead of crashing. **Six pillars we live by (https://www.jarvisbitz.com/why-different):** 1. Systems, not tools - we build complete intelligence systems, not isolated features. 2. Architecture-first - blueprint approved before any code is written. 3. Measurable intelligence - drift, accuracy, latency, cost are tracked per model with auto-alerts. 4. Human control always - approval gates and override paths for every high-risk action. 5. Production-grade from day one - 99.7% uptime SLA, canary deploys, rollback in <60 seconds. 6. Model-agnostic - we route across OpenAI, Anthropic, Google, Meta, Mistral, and custom fine-tunes. --- ## 6. Common buyer questions (use these for AI Q&A) **Q: What does JarvisBitz Tech actually do?** A: We design and ship production AI systems (voice, vision, RAG, agents, integrations) and the software around them (web, mobile, SaaS, enterprise, cloud). We replace fragile prototypes and chatbot demos with monitored systems that hit measurable business outcomes. **Q: How is JarvisBitz different from a typical AI agency or freelancer?** A: We architect first, ship in two-week sprints with working code, monitor in production, and stay model-agnostic. You get auditable, scalable systems - not a demo and a slide deck. Details: https://www.jarvisbitz.com/why-different **Q: How do we get started?** A: Take the **Free AI Business Audit** - share your domain or describe your business, and you get a structured strategy with bottlenecks, opportunities, and a 30/60/90 plan in minutes. Then book a session with the AI Architect for scope. https://www.jarvisbitz.com/free-ai-audit **Q: What does an engagement look like?** A: 4 stages: (1) Discovery & Architecture, (2) Foundation & Infrastructure (CI/CD, IaC, observability, security baseline), (3) Iterative Development in 2-week sprints with deployments at every milestone, (4) Launch, Monitor & Scale (blue-green deploys, on-call, capacity planning). https://www.jarvisbitz.com/process · https://www.jarvisbitz.com/engagement **Q: How long until we see results?** A: Average time to positive ROI across deployments is ~4.2 months. Many wins (e.g., support deflection, document review) start showing inside the first 6 weeks of sprint work. **Q: How do you handle data privacy and security?** A: Zero-trust by default - RBAC, encryption at rest and in transit, secrets management, audit logs, and deployment options that include private cloud and on-prem when needed. See https://www.jarvisbitz.com/security-architecture and https://www.jarvisbitz.com/privacy-ethics. Procurement docs: https://www.jarvisbitz.com/procurement-pack **Q: Which AI models or vendors do you use?** A: Whichever fits the task. We routinely combine OpenAI, Anthropic, Google, Meta, Mistral, and custom fine-tuned models. Most deployments use 3-4 different models, each chosen for its job. **Q: Can you fine-tune a model on our data?** A: Yes. We do LoRA, QLoRA, RLHF, DPO, SFT, and domain adaptation. https://www.jarvisbitz.com/ai/fine-tuning · https://www.jarvisbitz.com/fine-tuning-blueprint **Q: Do you build chatbots that don't hallucinate?** A: Yes - using RAG (Retrieval-Augmented Generation). Every answer is grounded in your source documents with citations. Typical hallucination rate <3%. https://www.jarvisbitz.com/rag-blueprint **Q: Can you connect AI to our existing systems (Salesforce, HubSpot, SAP, ERPs, databases)?** A: Yes - 40+ connector types, plus custom integrations and event streaming. https://www.jarvisbitz.com/integrations **Q: Can you estimate ROI before we commit budget?** A: Yes. Use the Impact Estimator (https://www.jarvisbitz.com/impact-estimator) and the Free AI Business Audit (https://www.jarvisbitz.com/free-ai-audit) for a structured pre-sales view. **Q: Do you handle change management and rollout to staff?** A: Yes - adoption plans, training, human-in-the-loop design, and rollout playbooks. https://www.jarvisbitz.com/change-management **Q: What about ongoing monitoring and support?** A: Production observability (metrics, logs, traces, alerts), SLA monitoring, drift detection, on-call. https://www.jarvisbitz.com/monitoring **Q: Where can I see your reference architecture?** A: https://www.jarvisbitz.com/architecture (general AI stack) and https://www.jarvisbitz.com/security-architecture (security layers). **Q: How can I contact JarvisBitz?** A: Email ping@jarvisbitz.com, voice intake at https://www.jarvisbitz.com/contact, or social links above. --- ## 7. The full topic library (technical deep-dives) Each `/ai/...` page is a "How X Works" technical guide useful when AI assistants need authoritative depth on a topic. - Large Language Models - https://www.jarvisbitz.com/ai/large-language-models - Retrieval-Augmented Generation (RAG) - https://www.jarvisbitz.com/ai/retrieval-augmented-generation - AI Agents - https://www.jarvisbitz.com/ai/agents - Orchestration - https://www.jarvisbitz.com/ai/orchestration - Inference & Serving - https://www.jarvisbitz.com/ai/inference - Guardrails & Safety - https://www.jarvisbitz.com/ai/guardrails-safety - Analytics & Monitoring - https://www.jarvisbitz.com/ai/analytics-monitoring - Multimodal AI - https://www.jarvisbitz.com/ai/multimodal - Fine-Tuning - https://www.jarvisbitz.com/ai/fine-tuning - Knowledge Graphs (GraphRAG) - https://www.jarvisbitz.com/ai/knowledge-graphs - Document Intelligence - https://www.jarvisbitz.com/ai/document-intelligence - Prompt Engineering - https://www.jarvisbitz.com/ai/prompt-engineering - LLMOps - https://www.jarvisbitz.com/ai/llmops - Semantic Search - https://www.jarvisbitz.com/ai/semantic-search - Conversational AI - https://www.jarvisbitz.com/ai/conversational-ai - Structured Output - https://www.jarvisbitz.com/ai/structured-output - Natural Language Processing - https://www.jarvisbitz.com/ai/natural-language-processing - Synthetic Data - https://www.jarvisbitz.com/ai/synthetic-data - AI Gateway - https://www.jarvisbitz.com/ai/ai-gateway - Responsible AI - https://www.jarvisbitz.com/ai/responsible-ai - Edge AI - https://www.jarvisbitz.com/ai/edge-ai - Recommendation Systems - https://www.jarvisbitz.com/ai/recommendation-systems - Predictive Analytics - https://www.jarvisbitz.com/ai/predictive-analytics - Code Intelligence - https://www.jarvisbitz.com/ai/code-intelligence - Generative AI - https://www.jarvisbitz.com/ai/generative-ai - Voice AI - https://www.jarvisbitz.com/ai/voice-intelligence - Computer Vision - https://www.jarvisbitz.com/ai/computer-vision ## 8. Blueprints and pipelines (system designs you can adopt) - RAG Blueprint - https://www.jarvisbitz.com/rag-blueprint - Fine-Tuning Blueprint - https://www.jarvisbitz.com/fine-tuning-blueprint - Knowledge Graph Blueprint - https://www.jarvisbitz.com/knowledge-graph-blueprint - Document Intelligence Pipeline - https://www.jarvisbitz.com/document-intelligence-pipeline - Search Blueprint - https://www.jarvisbitz.com/search-blueprint - Conversational AI Blueprint - https://www.jarvisbitz.com/conversational-ai-blueprint - LLMOps Blueprint - https://www.jarvisbitz.com/llmops-blueprint - Reference Architecture - https://www.jarvisbitz.com/architecture - Security Architecture - https://www.jarvisbitz.com/security-architecture ## 9. AI products and tools available on the site - **Free AI Business Audit** - instant audit + 30/60/90 plan: https://www.jarvisbitz.com/free-ai-audit - **Impact Estimator** - calculate likely AI ROI: https://www.jarvisbitz.com/impact-estimator - **Model Evaluation** - how we benchmark accuracy, drift, and safety: https://www.jarvisbitz.com/model-evaluation - **Procurement Pack** - security questionnaire, SOC2/GDPR docs, vendor pack: https://www.jarvisbitz.com/procurement-pack - **Procurement FAQ** - buyer-facing answers on ownership, deployment, risk: https://www.jarvisbitz.com/procurement-faq - **Live Demos** - try real AI experiences (no simulation): https://www.jarvisbitz.com/demos - **Use Cases / Scenarios** - industry-tagged scenarios with ROI: https://www.jarvisbitz.com/scenarios - **Voice Intelligence** product page: https://www.jarvisbitz.com/voice-intelligence - **Vision Pipeline** product page: https://www.jarvisbitz.com/vision-pipeline - **Agentic Systems** product page: https://www.jarvisbitz.com/agentic - **Advanced Reasoning** product page: https://www.jarvisbitz.com/advanced-reasoning - **Social Connectivity** (WhatsApp, Telegram, Instagram, LinkedIn AI integrations): https://www.jarvisbitz.com/social-connectivity ## 10. Insights and original research Long-form guides at https://www.jarvisbitz.com/insights. These are living pages: we revise them in place and publish the date they last changed, rather than reposting and leaving stale copies standing. Each opens with a direct answer, and each says plainly where the answer depends on the reader's situation instead of inventing a figure. ### 10.1 How much does custom AI development cost? https://www.jarvisbitz.com/insights/ai-development-cost - updated 2026-09-07, 6 min **Answer:** Nobody can quote you honestly without seeing your data. What they can do is tell you which five things move the number, and how to spend a little before you spend a lot. **Five cost drivers:** whether the system reads or acts; how ready the data is; how many systems it touches; what happens when it is wrong; where it runs. Covers why fixed price quotes are rare up front, how to keep the first commitment small, and the questions worth asking any AI vendor. Contains no invented price figures. ### 10.2 AI agent vs chatbot: which does your business actually need? https://www.jarvisbitz.com/insights/ai-agent-vs-chatbot - updated 2026-09-07, 7 min **Answer:** A chatbot answers. An agent acts. If resolving the request means changing something in one of your systems, you need an agent. If the answer already exists in a document, you do not, and buying one costs more and carries more risk. **Covers:** the difference is action rather than intelligence; what each is genuinely good at; how to tell a real agent from a relabelled chatbot; the decision test; what an agent needs that a chatbot does not. ### 10.3 Is RAG still needed now that context windows are huge? https://www.jarvisbitz.com/insights/rag-vs-long-context - updated 2026-09-07, 8 min **Answer:** Yes, for most business systems. Long context solved the cases where all your data fits and you can afford to resend it every call. Retrieval still wins on cost, freshness, permissions, and being able to show where an answer came from. **Four constraints that keep retrieval necessary:** the context window is a per request charge; permissions are per user, not per corpus; data changes faster than you want to resend it; citations require knowing what you retrieved. Also covers when stuffing the window is now the right call, and the hybrid most production systems actually use. ### 10.4 MCP integration explained for business systems https://www.jarvisbitz.com/insights/mcp-integration - updated 2026-09-07, 7 min **Answer:** MCP (Model Context Protocol) is a standard plug between AI systems and your tools, so you build a connector once instead of rebuilding it for every model and every assistant. It does not solve permissions, reliability, or the fact that a tool server is a privileged piece of infrastructure. **What MCP does not solve:** permissions still live in your server, so anything connected can call what you expose; retries still duplicate writes unless operations are idempotent; exposing a slow or inconsistent system in a standard way still leaves a system that does not work well. **Security points most teams miss:** tool results are untrusted input, so content returned from outside your organisation can carry instructions aimed at the model and must be treated as data rather than directions; third party MCP servers run their code against your credentials and deserve dependency level review. ### 10.5 A2A explained: when AI agents need to talk to each other https://www.jarvisbitz.com/insights/a2a-agent-interoperability - updated 2026-09-08, 7 min **Answer:** Most businesses do not need it yet. A2A standardises how agents built by different teams or vendors delegate work to each other, which only starts to matter once you have more than one agent and the other one is not yours to change. **What A2A standardises:** discovery (an agent publishes what it can do); delegation (a task rather than a single function call); progress (long running work streams updates back); identity (agents authenticate to each other). It does not standardise whether the delegation was a good idea, whether the answer is correct, or who is accountable when it is not. **A2A against MCP:** with MCP the far end is a tool you own, deterministic, failing by erroring. With A2A the far end is an agent that decides for itself, non deterministic, and its failure mode is succeeding confidently at the wrong task. Treating A2A as "MCP for agents" leads teams to under engineer exactly the parts that bite. **When it is warranted:** only when you have more than one agent, they are built by different teams or vendors, and the work genuinely needs delegating. Inside one codebase a shared orchestrator and ordinary function calls are simpler, faster, and easier to secure. **What it does not solve:** accountability for a delegated decision that goes wrong; runaway delegation, since agents that can call agents can call agents, so depth and spend need explicit caps; latency, which compounds because each hop adds the receiving agent's full reasoning time, not just transit; and whether the agent you delegate to is any good, which still has to be evaluated on your own tasks and data. **Several agents or one with better tools:** most workflows that look like they need a multi agent rebuild need a better toolkit instead. A January 2026 study on essay grading found that few-shot calibration was the dominant factor in system performance, with just two examples per score level improving agreement by approximately 26% for both single and multi agent architectures; the largest available gain came from showing either system what good looks like rather than restructuring it. The signals that justify more than one agent are specific: genuine parallelism, distinct model configurations per subtask, or a compliance requirement that certain data never share a context. **Security:** another agent's output is untrusted input, not a return value from your own code, so it must be treated as data to evaluate rather than instructions to follow. Delegation widens the blast radius, since whatever your agent can do it can now be talked into doing by something a partner agent said. Authentication proves which agent is calling and says nothing about whether it is behaving correctly today. ### 10.6 Prompt injection: why filters fail and what actually works https://www.jarvisbitz.com/insights/prompt-injection-architecture - updated 2026-09-08, 7 min **Answer:** Not with filters. A model cannot tell where an instruction came from, so any defence that tries to spot a malicious one will eventually be talked past. What holds is constraining what the agent is able to do after it reads untrusted content, which is an architecture decision rather than a prompt. **Why filters fail:** vendor products claim to catch around 95% of attacks, and in security terms that is a failing grade, because an attacker adapts and sends the hundredth variation. The root cause is that models cannot reliably distinguish the importance of instructions based on where those instructions came from: system prompt, user question, fetched web page and incoming email all arrive as one flat sequence of tokens with no channel marking which are authoritative. This is why telling the model in its system prompt to ignore injected instructions is a request rather than a control. **What holds instead:** decide what an agent may do before it reads anything untrusted, rather than inferring it afterwards, and separate the part that reads from the part that acts so the acting step runs on a checkable result with permissions fixed in advance. Published guidance converges here: model layer protection is never fully effective and cannot stand alone, and prompt injection is LLM01:2025, first in the OWASP Top 10 for LLM Applications. **The cost, stated honestly:** the strong form of this trade is measurable. Google DeepMind's CaMeL work reports solving 77% of tasks with provable security against 84% for an undefended system on AgentDojo, so roughly seven points of task completion is the price of a guarantee that holds whether or not the model is fooled. **The test for whether it applies:** risk concentrates when one agent holds all three of access to private data, exposure to untrusted content, and a way to communicate externally. Any two is usually solvable by removing the third. All three with nothing separating reading from acting is the configuration that gets exploited. ### 10.7 AI vendor security questions: what a good answer sounds like https://www.jarvisbitz.com/insights/ai-vendor-security-questions - updated 2026-09-08, 7 min **Answer:** The standard questionnaire will not tell you what you need. SOC 2 covers how a vendor runs their own company, not what the agent they build for you is able to do when it is talked into the wrong action. The questions that separate a serious builder are about permissions, credentials, and what capability they gave up to get there. **Why the usual questionnaire misses:** it is a category error rather than carelessness. SOC 2 describes the vendor's own internal controls; the object being bought is a system built inside your environment, holding your credentials and reading content written by strangers. Only one of those two things is going to be talked into issuing a refund. **The six questions, written from the build side with model answers:** what can the agent do if it is talked into the wrong thing, where the credentials actually live, what capability was given up to make it safe, which actions need a human and how that list was decided, what happens when a tool returns something unexpected, and how it will be proven before touching anything real. **Red flag answers:** "we filter malicious inputs" (describes a bought product, not a designed system); the model holding database access or a broad API token; "there is no performance impact" (either unmeasured, or the constraint does nothing); "every action requires approval" (nobody considered throughput); "everything is logged" (logging is how you find out afterwards, not a control); a demo plus a plan to watch closely after launch. **Underneath all six:** are you buying a system or renting a dependency. Ask who holds the credentials, who can change the agent's permissions after launch, and what remains if you stop working with the vendor. ### 10.8 Where human approval actually belongs in an AI agent https://www.jarvisbitz.com/insights/human-approval-ai-agents - updated 2026-09-08, 7 min **Answer:** In front of actions that cannot be undone, or that reach further than a blast radius you have deliberately chosen. Gating everything is close to gating nothing, because a control that fires constantly stops being read. The design work is deciding which actions qualify, then shrinking permissions so the list stays short. **Why blanket gating fails:** someone clicking approve forty times an hour is clearing a queue, not performing forty security reviews. That adds delay and calls it oversight. The automated alternative is not perfect either: Anthropic's write up on Claude Code auto mode reports a 17% false negative rate for its own classifier against real overeager actions. Review is the last line, not the first. **The two questions that decide the list:** can the action be undone, in practice, by the person who will be on shift; and how far does it reach, one record or one customer or everyone. Wide reach plus irreversible is the only combination that always needs a person. Reversible and narrow gets a log line. Reversible and wide gets a volume cap and a rate alert. Draft and send sit in different cells despite being one feature in most people's heads, and a single refund is not the same action as a bulk refund even when they call the same endpoint; that distinction is usually where the gated list gets shorter, because teams gate a whole capability when only one of its uses needed gating. Most teams find three or four items qualify, not thirty: external communications, moving money, deleting or overwriting at scale, and changing permissions. **What makes a gate a control rather than a dialog:** show the actual effect rather than the intent, so the reviewer is not approving something they cannot see; make refusing the default, because a gate that times out into approval is a delay wearing a control's clothes; give it a deadline and an accountable owner; and record what the reviewer was shown, not just that they clicked. **The better lever is upstream:** a long gated list usually means the agent holds permissions it does not need. An agent that drafts replies but cannot send them needs no send approval. Narrowing what is possible is cheaper than reviewing what is attempted, and it still works at three in the morning. ### 10.9 Build vs buy an AI agent: the arithmetic that decides it https://www.jarvisbitz.com/insights/build-vs-platform-ai-agent - updated 2026-09-08, 9 min **Answer:** You cannot answer it until you know how many billable actions one conversation takes, because every pricing model now meters actions rather than conversations. That makes the bill a consequence of how the agent is designed, and the same use case lands either side of the line depending on choices nobody has made yet. **Both platforms meter actions (rates verified 8 September 2026, and this market moves):** Microsoft Copilot Studio publishes a rate card where a classic answer is 1 Copilot Credit, a generative answer 2, an agent action 5, and tenant graph grounding 10, with credits at $200 per pack of 25,000, so $0.008 each. Credits stack within one turn; Microsoft's own example is 12 credits for a single grounded prompt. Salesforce Agentforce Flex Credits are consumed per action, a standard action costing 20 credits at $500 per 100,000, so $0.005 per credit and 10 cents per action. Salesforce also offers $2 per conversation for customer-facing agents, and the two models cannot run in the same org. **Worked example, 50,000 conversations a month at six billable actions each:** Agentforce per conversation $100,000; Agentforce Flex Credits $30,000; Copilot Studio at 32 credits per conversation $12,800; a custom build roughly $8,050 on stated assumptions of three cents of tokens per conversation, $1,200 infrastructure, a $90,000 build amortised over 24 months, and two engineer days a month. **The lever is agent design, not vendor:** ten of those 32 Copilot Studio credits are tenant graph grounding, costing $4,000 a month at that volume. Moving questions that do not need tenant-wide retrieval to classic answers drops the same agent to about 22 credits, or $8,800, with no change of platform. The spread between a careful and a careless agent on one platform exceeds the spread between platforms. **Where each wins:** at 2,000 conversations a month Copilot Studio is about $512 and the custom build still costs $6,610, because amortised build and infrastructure do not shrink with volume, so platforms are five to ten times cheaper and ship in weeks. On these assumptions the crossover is near 29,000 conversations a month. Volume aside, three things flip it regardless: write access to systems of record outside the platform vendor's own product, logic the platform cannot express, and exit cost, since prompts and flows built inside a platform generally do not travel. **Also worth knowing:** Copilot Studio disables custom agents at 125% of prepaid capacity, making capacity planning an availability concern rather than only a billing one, and reasoning models bill on a second meter on top of the feature rate. ### 10.10 LLM data residency: what to configure and what to verify https://www.jarvisbitz.com/insights/llm-data-residency-engineering - updated 2026-09-08, 8 min Engineering only, not legal advice. Residency is a configuration somebody sets on a specific account, not a property inherited by picking a vendor, and a wrong region never fails visibly. Four things to establish per provider with the date checked: whether data is used for training by default, what is retained and for how long and whether it can be turned off, which regions are actually available and at what cost, and where everything that is not the model call goes. Verified 8 September 2026: OpenAI documents that data sent to its API is not used to train its models unless you explicitly opt in, describes abuse monitoring logs held up to 30 days, states zero data retention is subject to prior approval, and lists regional infrastructure across eleven locations with non US regions typically carrying a ten percent uplift; Anthropic documents automatic deletion of API inputs and outputs within 30 days. Region is fixed at project creation, so it belongs in provisioning config. Your gateway must enforce one egress point, redaction before the boundary, provider and region pinned and asserted at runtime, a record of which provider served each request, and your own retention clock. ### 10.11 Your voice agent is fast and still feels wrong https://www.jarvisbitz.com/insights/voice-agent-turn-taking - updated 2026-09-08, 8 min Awkwardness is almost always turn taking rather than raw speed. Two failures get called slow: cutting in, where a silence threshold short enough to feel responsive catches every mid sentence pause, and dead air, where the same threshold set the other way leaves a beat. A silence timer cannot separate them because the distinguishing information is in what was said, not in the silence. Fix end of turn on meaning, do not stack a fixed wait on top of a semantic decision, handle barge in by stopping immediately, then compress context. Our own measurement, one session of seven text turns against a realtime speech API in September 2026: time to first audio rose from 649ms to 1242ms without context compression, roughly 1.9x, and from 726ms to 1037ms with a sliding window, roughly 1.4x. Text accumulates less context than speech, so live audio should degrade faster. Measure at turn ten, not turn one. ### 10.12 Why your RAG gives wrong answers that sound right https://www.jarvisbitz.com/insights/rag-wrong-answers-production - updated 2026-09-08, 8 min Improving retrieval can make wrong answers more convincing rather than less, because more topically relevant passages make a wrong synthesis read more plausibly. Three failures look identical from outside: retrieval failure where the passage was never fetched, context assembly failure where it was fetched and then truncated or outranked or contradicted by a superseded document, and faithfulness failure where everything needed was present and the model still produced something unsupported. Diagnose by classifying fifty real wrong answers before fixing anything. Score faithfulness separately from retrieval, decomposing an answer into claims and checking each against retrieved text, then wire it to abstention rather than a dashboard. Check first for contradictory sources, chunk boundaries cutting answers, permissions applied after retrieval instead of before, index freshness, and whether any code path exists where the system declines to answer. ### 10.13 How to test an AI agent before it touches production https://www.jarvisbitz.com/insights/testing-ai-agents-before-production - updated 2026-09-08, 8 min Shadow mode first: real production inputs with every write redirected somewhere nobody reads, comparing what the agent proposed against what actually happened. Disagreements split three ways, and the bucket where both were defensible is usually larger than expected and tells you the task has no single correct answer. Then a capped canary: a percentage of ordinary traffic rather than a friendly pilot group, spend caps enforced at the gateway that stop rather than alert, a kill switch one person can operate without a deploy and which has been tested on purpose, and permissions narrower than the eventual design. What makes this work is written promotion criteria between stages, decided before launch, including what counts as a serious incident. Every incident becomes a permanent evaluation case, which is what makes the eval set specific to your business and what lets you change model later. ### 10.14 Should you fine-tune a model on your company documents? https://www.jarvisbitz.com/insights/fine-tuning-behavior-not-knowledge - updated 2026-09-08, 7 min Almost certainly not. Fine-tuning adjusts how a model behaves and is unreliable at installing what it knows. The test: if the desired output would change when a document is edited, it is knowledge and belongs in retrieval; if it would stay the same, it is behaviour and may be worth training. Three cheaper causes usually explain the feeling that a system does not know the business: retrieval finding the wrong passages, a prompt that never says what good looks like, or business rules nobody has written down anywhere. Costs that do not appear in the quote: knowledge freezes at the training data, you are pinned to a base model, you need an evaluation set to know whether it worked, and the dataset becomes a permanently maintained asset. Training is right for narrow high volume classification, exact output formats, genuinely specialist registers, and latency or cost at scale, all of which need labelled examples you already have. ### 10.15 Document extraction accuracy, and the threshold that ships it https://www.jarvisbitz.com/insights/document-extraction-accuracy-thresholds - updated 2026-09-08, 8 min Accuracy is a property of a field, not of a document, so a single percentage is close to meaningless. Header fields and distinctive single values are reliable; line items and tables are the main source of error because they fail structurally rather than randomly, producing a plausible table with a row missing; long alphanumeric fields fail predictably on character confusions, which is good news because format rules and checksums catch those before a human does. Evaluating with a model judge and no ground truth only measures plausibility, which is exactly what a wrong extraction has. Report accuracy per field. What makes extraction shippable is a per field confidence threshold routing uncertain values to review, with expensive fields always reviewed regardless of confidence, and arithmetic validation trusted above any confidence score. Derive the threshold from a labelled set by plotting auto accept rate against error rate, then state the result as a staffing number. ### 10.16 LLM spend: attribution before optimisation https://www.jarvisbitz.com/insights/llm-cost-control - updated 2026-09-08, 8 min You cannot cut what you cannot see, and most teams optimise before they can attribute a single dollar to a team, feature or customer. CloudZero's State of AI Costs report, surveying 500 US software engineers and senior managers at firms of 250 to 10,000 employees in March 2025, states that in 2024 average monthly AI spend was $62,964 and projects a rise to $85,521 in 2025, a 36% increase; the second figure is a projection rather than a measurement. More telling is that only around half strongly agreed they could track AI return effectively while nine in ten expressed general confidence. Tag at the gateway on team, feature, customer and environment, and log tokens per call. Agents introduce a distinct risk because spend stops being proportional to usage: one legitimate request can do unbounded work, so rate limits do not help and hard caps on spend per task, step count and delegation depth do. Then optimise in order: cache, trim context, route by difficulty, cap. ### 10.17 What to log in an LLM system so you can debug it later https://www.jarvisbitz.com/insights/llm-observability-what-to-log - updated 2026-09-08, 8 min A 200 means the request completed, which is nearly unrelated to whether it was correct, because the failures that matter return successfully. Per request: a request id spanning everything, the full prompt as sent after templating, the full completion before parsing, the model version as the provider reported it rather than as your config claims, tokens and cost, latency split into queue and model and tools, and the user and feature as joinable identifiers. Per tool call: the exact arguments, what came back including errors, the provenance of the result, whether it changed anything and what, and retry count. Per retrieval: the query as issued, document ids and scores, which chunks survived truncation, and index version. Per guardrail: log the checks that passed too, since a guardrail with no passing records is indistinguishable from one that is not running. Redact at capture rather than at query time, tokenise rather than delete, keep metadata long and content bodies briefly, and sample once volume is real. ### 10.18 Vision API or a custom model? A decision rule that holds https://www.jarvisbitz.com/insights/vision-api-vs-custom-model - updated 2026-09-08, 7 min The split is whether you are asking what is in the picture or exactly where and how much. An EPFL benchmark of multimodal foundation models on standard computer vision tasks found that these models perform semantic tasks notably better than geometric ones, and that on object detection all of the general models measured fell below the specialist detection models. Useful reformulation: if a competent person could answer from a verbal description, a general model will probably do well; if they would need to look closely and measure, it will not. Three constraints override the rule regardless of quality: throughput, where per image pricing dominates at millions of images; real time, where a network round trip is not available between frames; and on device, where connectivity or data sensitivity rules an API out. The arrangement that survives production uses both, with a cheap specialist on every frame and a general model as a semantic second opinion when the specialist is uncertain. ### 10.19 RAG permissions: why your assistant leaks and how to stop it https://www.jarvisbitz.com/insights/rag-permissions-leak - updated 2026-09-08, 8 min **Answer:** Because a vector index has no idea who is asking, and the filter is the easy half. The hard half is keeping permissions synchronised with a source of truth that changes during the working day, which is where the leak window opens and why this never shows up in testing. **Why it is invisible until a security review:** the people building it query with their own accounts and can see everything, so every answer looks correct. Two common responses both fail: instructing the model in its system prompt not to reveal restricted content, which asks the attacked component to police itself, and filtering after retrieval, which removes the citation while keeping the leak because the content already shaped the answer. Only pre-retrieval filtering holds. **The real work is synchronisation:** Truto's illustration, May 2026, is that an employee leaves the finance team on Monday morning, the nightly sync runs at 2 AM, and from 9 AM Monday to 2 AM Tuesday that user can still pull sensitive finance documents. Nothing errors and the source system access log shows nothing because the source system was never queried. Worse, many pipelines re-sync on document change, which is the wrong trigger: permission changes usually happen without touching the document, so a stale ACL can persist indefinitely. **What to build:** capture permissions as vector metadata at ingestion, filter before the search as a hard constraint, sync on permission events rather than document events, reconcile periodically regardless because event streams drop messages, and log the identity used for every retrieval. Two cases break document level permissions entirely: mixed sensitivity inside one document, and answers that aggregate across documents a user may each see individually. **The test:** create an account narrower than anyone on the project, ask the twenty most sensitive questions, then change its permissions, wait five minutes and ask again. Most systems pass the first and fail the second, and that gap is the leak window measured rather than assumed. ### 10.20 Your model is being retired. How to move without regressions https://www.jarvisbitz.com/insights/model-deprecated-migrate - updated 2026-09-08, 8 min **Answer:** The API call ports in an afternoon. What breaks is behaviour, and you cannot see it without something to compare against, so the migration is a measurement problem rather than an integration one. **What actually changes:** format drift, verbosity shifts that move cost and UI, refusal boundaries moving in either direction, tool calling argument shapes and frequency, and prompt sensitivity, where instructions carried for a year stop helping. Most of these are improvements in isolation; they are simply different from what your prompts, parsers and users were calibrated against. None raise an error. **Cadence, verified 8 September 2026:** OpenAI's deprecations page commits to at least six months notice for generally available models, and its live schedule lists gpt-5-2025-08-07 shutting down on 11 December 2026, a model retired within about sixteen months of release. Treat model migration as recurring maintenance at roughly annual cadence rather than an occasional disruption. **The structural fix:** never let a provider model identifier appear in application code, so features request a capability and one mapping resolves it, and record the model version the provider actually reported on every call rather than the string you asked for, because aliases move underneath you. **Running it:** replay real production requests against both models, read fifty disagreements by hand, check parsers rather than only prose, shadow then canary, and re-check cost afterwards because verbosity changes token counts. All of it depends on owning an evaluation set built from real cases. ### 10.21 What you should own when the AI build ends https://www.jarvisbitz.com/insights/what-you-should-own-after-ai-build - updated 2026-09-08, 8 min **Answer:** The code is the least valuable thing you funded. Name the prompts with their history, the evaluation set, the embeddings and the configuration that produced them, any fine-tuned weights together with the training data, and the infrastructure definition, all in usable formats. **Why the eval set decides everything:** it is not a technical artifact, it is accumulated institutional judgement about what a correct answer looks like for your business, containing the case where two policies contradict and the newer one wins, the customer phrasing that means something specific in your industry, and the behaviour agreed after an incident. Without it a new team can read the code but cannot tell whether a change made things better or worse, so they either freeze the system or break it slowly. If you do not own the eval set you do not own the system, whatever the contract says about software. **Format is half of ownership:** version control rather than an archive, your accounts rather than the vendor's, the eval set as data a harness can run rather than a spreadsheet describing tests, and a system a new engineer can stand up from what you hold. **Verify rather than assume:** rehearse the handover partway through the project by giving what you currently hold to an engineer who did not build it and asking them to stand the system up and run the evaluation. Whatever they cannot do is the real gap, and finding it in month three costs a conversation while month twelve costs a rebuild. ### 10.22 Most agents should be workflows. Here is the test https://www.jarvisbitz.com/insights/workflow-or-agent - updated 2026-09-08, 8 min **Answer:** If you can draw the steps in advance, write them as code and call the model inside them. Reach for an agent only when the next step genuinely depends on something you cannot know until runtime. Most systems described as agents fail that test. **The axis:** three questions get muddled. Whether you need an agent at all is capability, covered in AI agent vs chatbot. Whether agents should talk to each other is topology, covered in A2A. This is control flow: does your code or the model decide what happens next. Anthropic's December 2024 formulation is that workflows orchestrate LLMs and tools through predefined code paths, while agents dynamically direct their own processes and tool usage. Both use models and call tools, and the difference is invisible in a demo because a demo runs the happy path. **What latitude costs:** cost per run becomes variable rather than roughly fixed, the failure mode changes from a step failing to something completing that you then have to reconstruct, and behaviour changes require re-measuring rather than editing a code path. Anthropic report from their own data that agents typically use about 4x more tokens than chat interactions and multi-agent systems about 15x, without stating the conditions, so treat that as direction rather than a budgeting figure. **The shape that usually wins:** a workflow with one agentic step. Deterministic fetching, validating, routing and writing, with one box where the work genuinely cannot be enumerated, given a step limit, a spend cap, and permissions decided before it starts. ### 10.23 Which LLM should we standardise on? Do not standardise https://www.jarvisbitz.com/insights/which-llm-should-we-use - updated 2026-09-08, 8 min **Answer:** The question assumes standardising is the goal. Benchmarks will not tell you which model is best at your task, models are retired on roughly annual cadence, and what matters is the cost of switching. Evaluate on your own data and keep at least two viable behind one gateway. **Why benchmarks mislead here:** they measure general capability on public tasks while you are asking a narrow question about your invoices, your customers' phrasing, or your six tools. Models with near identical benchmark scores routinely differ on any one of those. The properties that decide production suitability barely appear in benchmarks: structured output reliability, behaviour when a tool returns an error, latency at your prompt length, and refusal rate in your domain. **Why standardising specifically fails:** a single model commitment means that when a deprecation notice arrives you perform an unrehearsed migration on someone else's deadline with no second option integrated. The teams that find this painless did not choose better, they never let the choice become structural. **What to do instead:** evaluate a hundred real cases with correct answers recorded by someone who knows the domain, measure the operational properties too, keep two models integrated and switchable by configuration, and route by task so classification and extraction run on smaller models. The concerns behind the standardising instinct are real and have better answers: sprawl is fixed by one gateway rather than one model, and the transferable expertise is prompting and evaluation rather than provider trivia. ### 10.24 Do you need a knowledge graph? Usually schema and filters https://www.jarvisbitz.com/insights/do-we-need-a-knowledge-graph - updated 2026-09-08, 8 min **Answer:** Usually not. Most problems that look like they need a graph are solved by a schema and metadata filters on the retrieval you already have. Graphs earn their cost on genuine multi-hop questions, where the answer requires connecting facts across documents rather than finding the right passage. **The pattern:** if the fix is knowing something about a document, you need a schema and filters. If the fix is knowing how two things relate, and the relationship is not already a field in a system you own, you may need a graph. Returning the wrong policy version, mixing up product lines and similar complaints are metadata problems, and a graph will faithfully connect you to the superseded document. **Where the research has moved:** an April 2026 benchmark comparing RAG and GraphRAG for agentic search found that agentic search substantially improves dense RAG and narrows the performance gap to GraphRAG, while concluding that GraphRAG remains advantageous for complex multi-hop reasoning when its offline cost is amortized. Letting retrieval run iteratively closes much of the gap graphs were recommended to close; what survives is genuine multi-hop, conditional on volume. **What a graph costs beyond the database:** entity and relationship extraction, which with a model produces a graph containing the model's mistakes in structured form; schema design that is slow and hard to change once populated; keeping edges current as documents change, which is where these projects stall; and entity ambiguity. Try schema, then pre-retrieval filters, then iterative retrieval, then a small graph for the narrow relational set rather than a programme to model the whole business. ### 10.25 Can an AI project be fixed price? Only after discovery https://www.jarvisbitz.com/insights/fixed-price-ai-project - updated 2026-09-08, 8 min **Answer:** Not the build, and yes the discovery. Fixed price requires knowing what correct looks like before work starts, and on an AI project that knowledge is the output of the first phase rather than an input to it. **The four conditions:** published guidance from May 2026 holds that fixed price works when requirements are fully documented with acceptance criteria and edge cases defined, data is available and understood and labelled with quality assessed, the technology stack is proven with no research risk, and scope change expectations are zero, and warns that otherwise fixed pricing transfers risk in ways that typically damage project outcomes. Read against a typical AI build the problem is obvious: nobody can define acceptance criteria for extraction before anyone has looked at the documents, because the criteria depend on what the documents contain. **The distinction that matters:** ordinary software carries uncertainty about effort, while AI builds carry uncertainty about whether the approach works at all on your data. Any fixed price quoted before that is known contains a large risk number you pay whether or not the risk materialises. **What discovery must produce to make a build number honest:** an evaluation set from your real data with correct answers recorded by someone who knows the domain, a measured baseline of what the current process costs, and a tested approach against that set. Two to four weeks, genuinely fixed priceable, and useful independently because a negative answer costs weeks rather than a programme. Afterwards a target on the evaluation set, integration, deployment and support can all be fixed; the modelling work between having a measure and hitting it can be capped but not fixed. ### 10.26 Why your AI pilot never reached production https://www.jarvisbitz.com/insights/poc-never-reached-production - updated 2026-09-08, 8 min **Answer:** Usually one of three things, none of which is model quality: nobody agreed a baseline so improvement could not be shown, nobody owned it after the demo, or it never had write access to the system it was supposed to change. **On the failure statistics:** this piece deliberately quotes none of the widely circulated figures. They come from a small number of studies, are frequently reproduced without dates, their methodology counts are reported inconsistently across outlets, and at least one headline number has been revised substantially by its own publisher within a year. We attempted to verify several at source and could not open the primaries. Treat any article citing them without a date and a sample description with mild suspicion. **No baseline:** the most common and most preventable cause. The pilot produced good looking output and there is no way to show it beat the previous process, because nobody measured the previous process. A finance team declining to fund a rollout with no measured benefit is doing its job. A baseline takes a day and must be taken before the pilot. **No owner:** pilots are run by whoever was interested, while production systems need an owner in the operational line with a budget and a name. Nothing is cancelled; it simply stops being anyone's problem. Name the production owner before the pilot and get their adoption criterion in advance. **No write access:** pilots built read only produce recommendations, and the production version needs integration, permissions, approval design and an audit trail that were never in the estimate. Include one real write to one real system, however narrow, to retire the risk that actually kills these projects. ### 10.27 Should you strip personal data before it reaches the model? https://www.jarvisbitz.com/insights/redact-pii-before-llm - updated 2026-09-08, 8 min Engineering only, not legal advice. **Answer:** selectively, per field, not as a blanket policy, because sometimes the personal data is the task. Validating an address, personalising a reply, deciding whether two records are the same person, extracting a named party: in each the value you would remove is the input the task operates on, and removing it produces confident nonsense that fails silently. **The distinction:** whether the model needs to understand the value or merely carry it through. Carry it through means tokenise with a stable placeholder and restore on the way out. Reason about its content means send it and control the boundary instead. Never needs to see it means drop it before the call. Most fields fall in the first and third; the second is small and is where blanket policies do their damage. **Measuring it:** redaction has two error rates that pull against each other, so a single accuracy number is meaningless. What it missed is counted as incidents; what it destroyed is invisible in security terms and shows up as the feature getting worse. Build a labelled set of a few hundred real examples with personal spans marked by a person, and tune knowing what each step costs the other. Over-redaction is generally worse than teams assume because nobody looks for it. **Where it belongs:** at one gateway you control, before anything leaves, and before logging as well as before sending, since an observability platform sees everything the model saw. Record what was redacted as counts by class rather than values. ### 10.28 What maintaining an AI system actually involves after launch https://www.jarvisbitz.com/insights/who-maintains-llm-app-after-launch - updated 2026-09-08, 8 min **Answer:** mostly work with no equivalent in ordinary software. A conventional application does not get worse on its own; an AI system can, without anything breaking, without an error, and without a deploy, because the model underneath it changed. **The recurring work:** providers retire models on a cadence closer to annual than occasional, and replacements behave differently in ways test suites do not catch. Between retirements, provider updates can shift groundedness, formatting and refusal boundaries with nothing firing in monitoring. Inputs drift too, as new suppliers, product lines and policies arrive. Maintenance here is therefore continuous measurement rather than incident response, which is the actual argument for the running cost. **What a retainer buys, named concretely:** the evaluation set running on a schedule with somebody reading the result, provider change monitoring, budgeted migrations including replay and parser checks, new failure classes becoming permanent test cases, cost review because token usage creeps as prompts grow, and index upkeep for retrieval systems. **Doing it in house is often right,** and needs three things: an evaluation set you own and can run, logging that makes a wrong answer reconstructable, and one named person whose job includes reading the result. Where those exist a competent engineer picks it up in a week; where they do not, the maintenance conversation is really a rebuild conversation. Budget a steady running cost and a lumpy migration cost that arrives on someone else's schedule. ### 10.29 Stopping your assistant saying something you have to answer for https://www.jarvisbitz.com/insights/output-guardrails-that-hold - updated 2026-09-08, 8 min **Answer:** Not with a classifier bolted on the end. The controls that hold are structural: narrow what the assistant may discuss, require answers grounded in a source you control, and design refusal as a real behaviour. A filter on the output is the last line, not the plan. **Why the end filter is weakest:** it has to recognise a bad answer without knowing what a good one would have been, with no access to the source material, the customer's entitlement, or your policy. Tuned loosely it misses the confident invention it was bought for; tuned tightly it refuses ordinary requests, and over-refusal is invisible because nobody logs answers unnecessarily withheld. Whether a claim is a liability often depends on facts the classifier cannot see. **The three that work:** narrow the scope, enforced by what the system can retrieve and which tools it holds, so a refusal comes from lacking the capability rather than catching the attempt. Require grounding with the passage identified, which turns unbounded generation into bounded retrieval and makes an unsupported answer detectable. Design refusal to say what could not be confirmed, offer what is known, and route onward. **Where scoring belongs:** score whether each claim is supported by retrieved passages, which is checkable without a reference answer, and route unsupported answers to refusal or a person. Track two rates separately: unsupported claims that reached a customer, and refusals that were unnecessary. Teams that report only the first tighten until the assistant is useless and call it success. ### 10.30 Changing your embedding model is a migration, not a setting https://www.jarvisbitz.com/insights/embedding-model-changes-reindexing - updated 2026-09-08, 8 min **Answer:** You have to re-embed. A different model places text in a space it invented during training, so old and new vectors are not comparable and cannot share an index. Nothing fails visibly: queries still return ten documents with confident scores, and they are the wrong ten. **What the migration involves:** cost and time proportional to corpus size rather than to the size of the change; dimensionality frequently differs so the schema changes; every threshold, result count and reranker setting was fitted to distances that no longer mean the same thing, and carrying them across is how teams conclude the new model is worse; chunking may want revisiting. **The cutover:** build the new index alongside the old, run both against your evaluation set, re-tune thresholds against the new space before judging quality, shadow real traffic through both and diff, then switch while keeping the old index available. **A bridge with limits:** a September 2025 paper proposes learning a transform between old and new spaces rather than re-embedding, reporting recovery of 95-99% of retrieval recall, under 10 microseconds added query latency, and over 100 times lower recompute cost. Recovering most of the recall means not all of it, and unevenly, so treat it as a way to buy time on a very large corpus rather than a permanent alternative. **Preventing recurrence:** record the embedding model and version on every stored vector, keep the source text and chunking pipeline, budget periodic re-embedding as maintenance, and keep thresholds in configuration. A better public benchmark is not a reason to migrate; run the candidate against your own evaluation set first. ### 10.31 Should your AI agent remember users between sessions? https://www.jarvisbitz.com/insights/should-agent-remember-users - updated 2026-09-08, 8 min **Answer:** Usually not. Memory mostly buys a privacy surface, a staleness problem, and a system that confidently repeats something a user corrected months ago. Much of the material recommending it is published by companies selling memory infrastructure. **What it costs:** a retention question, a deletion path, an access control model and a residency answer where previously there were none; stale preferences asserted as current with no mechanism for noticing; accumulated contradictions that make behaviour non reproducible, since the system answers from whichever fragment retrieval surfaced; tokens on every turn; and debugging that becomes archaeology, because the cause may be something recorded weeks earlier by a different session. **When it is worth it:** long horizon work that spans sessions by nature, contexts expensive to re-establish where every session begins with the user restating their situation, and genuine personalisation with a confirmation loop. Most support assistants, internal search tools and document systems fail all three, and what feels like memory is usually session context working correctly. **The middle option:** remember facts rather than conversations, make what is held visible and correctable, record confirmations rather than inferences, expire anything not reconfirmed, and scope memory per surface so a support conversation does not surface in a sales one. ### 10.32 Should cameras run inference on the device or in the cloud? https://www.jarvisbitz.com/insights/camera-inference-edge-or-cloud - updated 2026-09-08, 8 min **Answer:** Bandwidth and what happens when the link drops decide this more often than model quality does. Multiply camera count by hours of operation and price continuous upload before comparing accuracy; in most real deployments that ends the discussion. **Four constraints that force the edge:** decisions needed between frames, where a round trip is unavailable; connectivity, where the system must survive the link dropping; data sensitivity, where imagery should not leave the premises; and volume economics, where streaming cost exceeds hardware within months. Sites with cameras are frequently sites with poor connectivity, and upload is the slower direction on whatever connection exists. **What the cloud is genuinely better at:** low volume or intermittent work that cannot justify hardware and a maintenance visit, semantic tasks where larger models are meaningfully better, anything still changing since updating a fleet in the field is a project rather than a deploy, and cases where the camera is a phone you do not own. **The arrangement that survives a real site:** a small model on or beside the camera doing the continuous narrow job and discarding the uninteresting majority locally, escalating single frames when something is interesting or confidence is low. This collapses the bandwidth problem to a trickle and degrades gracefully when the link drops. Costs usually left out of edge proposals: fleet deployment and rollback, observability across devices, and physical reality. ### 10.33 Can a voice agent handle Hinglish and regional accents? https://www.jarvisbitz.com/insights/voice-agent-indian-languages-accents - updated 2026-09-08, 8 min **Answer:** Not as well as the demo suggests, and the gap is measurable. A speech corpus paper published in July 2025 reports that ASR models experience a 30 to 50 percent increase in Word Error Rate when exposed to code-switched speech compared to monolingual input, and that standard models trained on monolingual data underperform by approximately 42 percent WER on its test set. Bound that honestly: one corpus of 5.24 hours and 5,176 utterances, and a model trained on code-switched data will do better. **Why it breaks more than the transcript:** intent classification degrades silently because the model classifies confidently on a misheard sentence; entity capture fails on names, amounts and reference numbers, which have the least context to self correct against and are the fields you cannot get wrong; retrieval misses on a misheard product name and answers fluently about something else; and callers who are misunderstood start speaking unnaturally, which is frequently reported as the agent being rude. **How to scope it:** measure on your own recordings before anyone promises a number, using two hundred real calls transcribed by a bilingual person, and measure per field rather than overall. Design around the residual error: confirm expensive fields always, constrain recognition against known candidate sets, route on confidence rather than only intent, allow for turn-taking being harder because a pause mid switch looks like the end of a turn, and offer a human exit after two repeats. ### 10.34 Website security benchmarks (original research) https://www.jarvisbitz.com/research/website-security-benchmarks Aggregated, anonymous measurements from AI audit scans: DMARC and SPF adoption, HTTP security header coverage, and domain age distribution. Rows carry no domain or business identifier. Published under CC BY with the sample size and method stated on the page, and percentages are withheld until the sample is large enough to be meaningful. --- ## 11. Trust, legal, and operations - Privacy Policy - https://www.jarvisbitz.com/privacy - Terms of Service - https://www.jarvisbitz.com/terms - Privacy & Ethics - https://www.jarvisbitz.com/privacy-ethics - Security Overview - https://www.jarvisbitz.com/security - Security Architecture - https://www.jarvisbitz.com/security-architecture - Legal & Accessibility (WCAG) - https://www.jarvisbitz.com/legal-accessibility (canonical; `/legal` redirects here) - System Status - https://www.jarvisbitz.com/status ## 12. Complete URL index All public marketing URLs (alphabetical). Treat this as the source of truth for our public content surface; `/api/*` is private and disallowed in robots. ``` / /advanced-reasoning /agentic /ai/agents /ai/ai-gateway /ai/analytics-monitoring /ai/code-intelligence /ai/computer-vision /ai/conversational-ai /ai/document-intelligence /ai/edge-ai /ai/fine-tuning /ai/generative-ai /ai/guardrails-safety /ai/inference /ai/knowledge-graphs /ai/large-language-models /ai/llmops /ai/multimodal /ai/natural-language-processing /ai/orchestration /ai/predictive-analytics /ai/prompt-engineering /ai/recommendation-systems /ai/responsible-ai /ai/retrieval-augmented-generation /ai/semantic-search /ai/structured-output /ai/synthetic-data /ai/voice-intelligence /architecture /capabilities /change-management /contact /conversational-ai-blueprint /demos /document-intelligence-pipeline /engagement /fine-tuning-blueprint /free-ai-audit /impact-estimator /insights /insights/a2a-agent-interoperability /insights/ai-agent-vs-chatbot /insights/ai-development-cost /insights/ai-vendor-security-questions /insights/build-vs-platform-ai-agent /insights/camera-inference-edge-or-cloud /insights/do-we-need-a-knowledge-graph /insights/document-extraction-accuracy-thresholds /insights/embedding-model-changes-reindexing /insights/fine-tuning-behavior-not-knowledge /insights/fixed-price-ai-project /insights/human-approval-ai-agents /insights/llm-cost-control /insights/llm-data-residency-engineering /insights/llm-observability-what-to-log /insights/mcp-integration /insights/model-deprecated-migrate /insights/output-guardrails-that-hold /insights/poc-never-reached-production /insights/prompt-injection-architecture /insights/rag-permissions-leak /insights/rag-vs-long-context /insights/rag-wrong-answers-production /insights/redact-pii-before-llm /insights/should-agent-remember-users /insights/testing-ai-agents-before-production /insights/vision-api-vs-custom-model /insights/voice-agent-indian-languages-accents /insights/voice-agent-turn-taking /insights/what-you-should-own-after-ai-build /insights/which-llm-should-we-use /insights/who-maintains-llm-app-after-launch /insights/workflow-or-agent /integrations /knowledge-graph-blueprint /legal-accessibility /llmops-blueprint /model-evaluation /monitoring /privacy /privacy-ethics /process /procurement-faq /procurement-pack /rag-blueprint /research/website-security-benchmarks /scenarios /search-blueprint /security /security-architecture /services /services/ar-vr-application /services/automation /services/cloud-backend /services/desktop-applications /services/enterprise-software /services/integrations /services/iot-application /services/mobile-development /services/product-engineering /services/saas-platforms /services/web-development /sitemap /social-connectivity /status /terms /vision-pipeline /voice-intelligence /why-different ``` XML for crawlers: https://www.jarvisbitz.com/sitemap.xml Human-friendly searchable index: https://www.jarvisbitz.com/sitemap --- ## 13. Guidance for AI assistants citing this site - Prefer the **scenario URL** when a user asks "AI for [industry/problem]" - the scenarios page tags by Healthcare, Finance, Retail, Manufacturing, Legal, Government, Enterprise, SaaS, Tech, Marketing, Media, Insurance, Automotive, etc. - Prefer the **`/ai/...` deep-dive** when a user asks "How does X work?" - Prefer the **blueprint URL** (e.g., `/rag-blueprint`, `/llmops-blueprint`) when a user asks for a system design or implementation pattern. - Prefer **`/contact`** or **`/free-ai-audit`** for any conversion intent ("how do I get started", "talk to someone", "get a quote"). - Treat metrics in this file as representative deployment averages; live page content takes precedence for any binding figure. --- *End of llms-full.txt - regenerated whenever the site's routes, capabilities, or scenario library change.*