Grimvane AI

Intelligence,
engineered.

Inference engines, autonomous agents, knowledge retrieval, coding tools, and the infrastructure underneath them. Built from scratch. Local-first.

What we build

AI that works
for you.

Enginebuilding

Crucible

From-scratch LLM inference engine. Direct GGUF parsing, full transformer forward pass, GPU compute via CUDA. No third-party wrappers, no framework dependencies.

Inference Engine From Scratch Python CUDA
Explore Crucible

Productbuilding

Shoal

Multi-dock AI coding platform. Connects local inference, frontier APIs, and CLI tools through a unified architecture. No telemetry, no cloud requirement.

Code Assistant Multi-dock Local-first Go
Explore Shoal

Productbuilding

Hearth

Self-hosted conversational AI with document-grounded memory. Chat across backends with model switching and session persistence, or index your documents and codebases and get answers cited back to the source. Retrieval is opt-in.

Chat Interface Retrieval Self-hosted Multi-backend Python
Explore Hearth

Library Suitegrowing

Flower Garden

Six zero-dependency libraries for AI infrastructure. Full-text search, vector storage, graph databases, tiered caching, encrypted storage, and structured data. Pluggable backends throughout.

Aster Camellia Dahlia Lotus Thistle Wisteria Python GDScript
Explore Garden

Notation

Agents do the reading now.
It's a given.

Every format they read was built for someone else. YAML, JSON and TOML all assume a person is doing the reading, which was true when they were written.

GVN is GrimVane Notation, written for the reader you actually have. It puts one fact on a line and gives that line its own address, so nothing it says depends on the lines above it. Ordering is canonical, which means the same facts always come out as the same bytes.

We run on it. Our workspace graph and project maps are written in GVN, and the whole specification fits on one page.

# one fact per line, addressed at the line
hearth.class = "self-hosted-chat"
hearth.retrieval = true
hearth.tech[0] = "python"
# relationships are statements too
rel.hearth.camellia.type = "consumes"
rel.corvath.wisteria.type = "consumes"
# sorted, deduplicated, byte-stable

It reads as accurately as JSON and costs about what YAML costs. It needs half the lines, and it invents fewer answers than either.

The same graph, read by agents GVN YAML JSON
Fields recalled correctly 100%96.7%100%
Tokens spent, against YAML +2.9%baseline+20.4%
Answers invented rather than refused 61211
Lines to carry the graph 1,1851,9522,841

Measured across 365 graded agent runs.

Engine

Built from scratch.

Crucible implements a full transformer forward pass with direct GGUF parsing, RoPE embeddings, grouped-query attention, and GPU compute. No wrappers. No framework dependencies. Every layer is ours.

Direct GGUF Parsing

Memory-mapped tensor access with dequantization for F16, Q4_0, Q4_K, Q6_K, and Q8_0.

Full Forward Pass

RMSNorm, RoPE, grouped-query attention, SwiGLU FFN. The entire transformer stack, implemented from first principles.

GPU Compute

CUDA acceleration via CuPy with streaming token output and nucleus sampling.

Privacy

Private by default.

Every product in the Grimvane AI ecosystem runs locally. No API keys required, no cloud round-trips, no telemetry. Your data never leaves your machine unless you choose otherwise.

Zero Telemetry

No usage tracking, no analytics callbacks, no phone-home behavior. Period.

Hardware Detection

Automatic platform and GPU detection. Metal, CUDA, ROCm, or CPU fallback.

Offline-capable

Hearth and Crucible run entirely offline once models are downloaded.

Architecture

Composable by design.

Flower Garden provides six standalone libraries for search, vectors, graphs, caching, encryption, and structured data. Each one is zero-dependency at its core with pluggable storage backends.

Pluggable Backends

In-memory, SQLite, and PostgreSQL backends for every library. Swap without changing application code.

Domain-agnostic

Built for AI infrastructure but designed to work anywhere. No opinions on your architecture.

Interoperable

Hearth uses Camellia for vectors. Corvath uses pluggable memory. The pieces compose naturally.

Capabilities

What we do.

Core competencies across the AI stack, from low-level model work to high-level agent orchestration.

01

From-Scratch Inference

Direct GGUF parsing, full transformer forward pass, and GPU-accelerated token generation with no framework dependencies.

02

Autonomous Agents

Multi-step reasoning with task classification, tool orchestration, role-based security, and automatic rollback.

03

Knowledge Retrieval

Document and codebase indexing with vector search, source attribution, and fully offline question answering.

04

Code Generation

Multi-dock architecture connecting local inference, frontier APIs, and CLI tools through a unified coding platform.

05

Conversational AI

Backend-agnostic chat with model switching, session persistence, and automatic hardware detection.

06

AI Infrastructure

Full-text search, vector storage, graph databases, tiered caching, encrypted storage, and structured data libraries.

Get started

Ready to build
with AI?

Whether you need an inference engine, an autonomous agent, or local AI infrastructure, we'd like to hear about it.