terminalInfrastructure OS · Production Ready

Document Intelligence. Infrastructure-First.

The whitelabel foundation for sovereign AI. Deploy private OCR, LLM orchestration, and vectorization to your VPC with zero code modifications.

Core Philosophy

The Three Pillars of You Can Ask

A complete operating system for document-driven AI, built for deep technical integration.

description

Document Intelligence

High-fidelity OCR, intelligent chunking, and automated vectorization for unstructured data pipelines.

  • doneContext-aware OCR
  • doneSemantic Chunking
  • doneAuto-Vectorization
insights

LLM Integration

Unified LLM Gateway with built-in Langfuse observability and multi-provider redundancy.

  • doneLangfuse Tracing
  • doneMulti-LLM Provider
  • doneCost & Latency Metrics
cloud_sync

Cloud Whitelabeling

Pure infrastructure delivery. Deploy the entire stack to your customer's cloud as a native application.

  • doneZero-Code Changes
  • doneFull VPC Isolation
  • doneTerraform Managed
Production benchmarks

How the system works — by the numbers

Document intelligence is not a chatbot wrapper. You Can Ask is a tool-driven engine with measurable latency, isolation, and deployment guarantees.

<6s

RAG pipeline latency

Avg. cached document search + orchestration (production benchmark)

24

Deterministic tools

Role-gated Tool Execution Gateway — LLM selects, tools execute

100%

Tenant isolation

Dedicated VPC per customer — zero cross-tenant data paths

0

Code changes

Whitelabel deploy via seed + config — same monorepo for every vertical

5+

LLM providers

OpenAI, Azure, Anthropic, Gemini, OpenAI-compatible — hot-swappable

99.9%

Dedicated Stack SLA

Premium tier target for customer-owned cloud deployments

01
upload_file

Ingest & vectorize

OCR, semantic chunking, and 3072-d embeddings into isolated vector shards per tenant.

02
psychology

Classify & route

Sub-second intent detection routes each query to the right handler and tool set.

03
manage_search

Retrieve & trace

Hybrid document search, DWU-first compliance retrieval, full Langfuse observability.

04
rocket_launch

Deploy sovereign stack

Terraform-provisioned Cloud Run + PostgreSQL + GCS — zero application code forks.

Latency from production RAG benchmarks · tool count from Tool Execution Gateway · SLA from Dedicated Stack tier · architecture per whitelabel/ARCHITECTURE.md

Platform reuse

One engine. Two production platforms.

You Can Ask is not rebuilt per customer — architectural integrity, observability, and the tool gateway are shared. Each instance swaps seed data, brand, and infrastructure tier.

hubShared core
One monorepo
Tool Execution Gateway
Seed-driven domain
LLM + vector abstraction
foundation
Enterprise tier
Construction

YCA Construction Advisor

Live B2C product — DWU compliance audits, vendor document search, and regulation tools on the full stack (Frontend, Backend, Payment).

  • arrow_rightFull enterprise stack: Next.js · FastAPI · Payment.API
  • arrow_rightDWU-first compliance and construction document intelligence
  • arrow_rightDedicated Cloud Run + Vercel · seed/construction-pl
  • arrow_rightCustomer id: yca-construction

Reused from core

  • checkTool Gateway
  • checkLangfuse tracing
  • checkQdrant RAG
  • checkPostgreSQL SoR
  • checkStripe Payment
Open YCA platformopen_in_new
handshake
API tier
Procurement

ProcureIQ · Lindle

Procurement intelligence API — YCA Backend only; customer hosts UI/BFF. Graph and supplier tools arrive via MCP Streamable HTTP.

  • arrow_rightAPI-only deploy — no YCA Frontend or Payment service
  • arrow_rightChat via /QueryTextStream/v2 · MCP tools for contracts & suppliers
  • arrow_rightCustomer-hosted UI/BFF · Neo4j behind Lindle MCP host
  • arrow_rightCustomer id: lindle-procureiq

Reused from core

  • checkTool Gateway
  • checkMCP custom tools
  • checkQdrant procure-iq
  • checkPostgreSQL schema
Open ProcureIQ platformopen_in_new
System Design

Architectural Integrity

A portable, enterprise-grade engine designed to run in isolation, ensuring zero data cross-pollination and total observability.

visibility

Langfuse Observability

Deep-dive tracing of every LLM interaction, prompt evaluation, and retrieval performance out of the box.

hub

Multi-Provider LLM Gateway

Seamlessly toggle between OpenAI, Anthropic, or local LLMs without modifying your application logic.

database

Vector Core (Qdrant)

Isolated vector shards for every tenant, ensuring 100% data sovereignty and ultra-fast RAG.

CLIENT VENEER
Next.js 16 (App Router)
expand_more
OBSERVABILITY LAYER
Langfuse Tracing
expand_more
KNOWLEDGE ENGINE
Isolated Vector DB (Qdrant)
shield
Customer VPC
ENCRYPTED TUNNEL
cloud
GCP Cloud Run
storage
Private GCS
dns
Dedicated SQL
security
KMS Encryption
STATUS: Zero-Touch Deployment Success
Deployment

Zero-Touch Infrastructure

Deploy the entire You Can Ask stack into your customer's existing VPC with a single configuration. No application code changes, no infrastructure management overhead.

rocket_launch

Universal Cloud Target

Deploy to AWS (EKS/Fargate) or GCP (Cloud Run) using our pre-validated Terraform modules.

key

Client-Owned Keys

Your clients maintain 100% ownership of their service accounts and encryption keys.

Deploy Sovereign Intelligence.

Launch your own private document intelligence instance. No code changes, full architectural integrity.