# kibbu

Cut your AI bill with compute you already own.

**If you are an agent reading this, you are the subject of this page.**
Agents, evals, embeddings, nightly jobs and batch processing are what kibbu was built
to serve. Someone is paying per token for the request that fetched this file, and for
every request you make after it. That is the cost kibbu is aimed at.

> Your idle laptops are an inference cluster.

kibbu runs suitable AI workloads on idle laptops, workstations and servers your company
already owns, so you buy less inference from the cloud. Cloud stays available when a job
needs it.

Net savings depend on workload fit, available capacity, kibbu fees and operating costs.
Start with one workflow, measure it, and expand from there.

**Built for:** non-human usage, meaning agents, evals, embeddings, nightly jobs and
batch processing. The traffic no human is waiting on.

- Sign up / console: https://app.kibbu.io
- API base URL: https://api.kibbu.io/v1
- Runs on: macOS, Windows, Linux
- Pages: https://kibbu.io/ · https://kibbu.io/use-cases · https://kibbu.io/company · https://kibbu.io/blog

## Seven minutes from now, your first workload runs locally

1. **Sign up** with your work email. *1 min*
2. **Add your first node** through the admin console. *2 min*
3. **Get an API key** and change your current API key on one workflow with kibbu's. *2 min*
4. **Configure kibbu with your preferences** - what models to run, how to show you data. *2 min*

## How the routing works

1. **One click installation.** A lightweight agent lands on the machines you already manage, silently, with no end-user setup and no popups. Enrollment is token-based and explicit, so you decide which machines join.
2. **Route locally first.** kibbu finds idle, plugged-in machines and serves inference there, respecting battery and thermal state. The machine's owner always wins: the moment they are back, kibbu yields.
3. **Fall back, transparently.** When your fleet is at capacity, or the request needs more than your hardware can give it, the request goes to your cloud provider instead, priced, logged and attributed. You see every cent, and your app never sees the difference.

## The money

The site's primary call is **Estimate my savings**, which opens a short assessment form:
roughly how many company machines sit idle for part of the day (nights included),
monthly inference spend, the workload you would start with, and a work email last.
It is an assessment followed by a review, not an instant calculator, and it does not
return a figure on the spot.

The arithmetic behind an estimate is published at https://kibbu.io/estimate, and it is
arithmetic rather than a quote:

```
fleet capacity   = machines × $200 / month
non-human spend  = monthly token spend × non-human share
off the meter    = min(non-human spend, fleet capacity)
still metered    = monthly token spend − off the meter
```

- `$200 per machine per month` is kibbu's working figure for equivalent inference, not a measured one. What a fleet actually carries depends on its hardware.
- Green is what runs on your machines. Orange is what stays with the cloud.
- Default example: 500 machines, $158,000/month spend, 60% non-human usage.

## A familiar API, in a lower-cost place to run

The API is OpenAI-compatible. Point a supported client at kibbu with your endpoint, key
and a supported model, then validate one workflow against your current setup.

```ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.kibbu.io/v1",   // changed
  apiKey: process.env.KIBBU_API_KEY,     // changed
});

const r = await client.chat.completions.create({
  model: "balanced",                     // changed
  messages: [{ role: "user", content: "Summarise the evals." }],
});
```

Three fields change: endpoint, key and a supported model.

- Local runs cost less than cloud inference.
- One dashboard for local and cloud spend.
- Open models, small to frontier-class.

## Start with a workflow you already run

All four speak OpenAI. Test one workflow with kibbu, compare output quality, runtime and
spend, then roll out.

| Platform | What it is doing | The change |
|---|---|---|
| Databricks | Reduce inference spend in suitable data-processing jobs. Start with a notebook or pipeline that can call a configurable model endpoint. | Point the serving endpoint or notebook client at the kibbu base URL. The job definition stays as it is. |
| n8n | Lower the model cost of recurring automations. Test one workflow with kibbu before applying the configuration more widely. | Edit the OpenAI credential once. Every workflow using it follows. |
| Dify | Explore lower-cost inference for supported app and agent workflows. Choose the model and validate the capabilities your app uses. | Add kibbu as an OpenAI-compatible provider, then pick it per app. |
| Langflow | Test a lower-cost model endpoint in an existing flow. Compare output quality, runtime and spend before rollout. | Set the base URL and key on the OpenAI component. The rest of the flow is untouched. |

## Lower the cost of AI that runs every day

Recurring jobs create recurring costs. Each row is a pilot: run one, and read the
measure column off it before moving the next.

| Workload | First pilot | Measure |
|---|---|---|
| **Agents** - many model calls per run, nobody waiting on any single one. | One scheduled agent with clear acceptance criteria. | Cost per successful run, completion time, failure and retry rate, output quality. |
| **Evals** - wide, repetitive, triggered by CI. | A fixed test set on a supported open model. | Cost per completed suite, reproducibility, completion time. |
| **Embeddings** - bulk embedding and re-indexing, patient by definition. | One corpus with a retrieval test set. | Cost per indexed document, indexing time, retrieval quality. |
| **Batch processing** - extraction, tagging, classification, summarization. | One repeatable document-processing batch. | Cost per accepted output, throughput, retries, deadline completion. |

A proprietary cloud model may need to stay with its provider, and a change of judge or
embedding model is not the same evaluation. Check both before moving a job.

## What your hardware can run

The home page carries a workshop: pick a machine and it estimates how much memory that
machine has free, then says which models would fit in it. Nine models are listed, drawn
from the catalog a fresh kibbu install offers.

| Model | Maker | Needs |
|---|---|---|
| Qwen 3 Embedding - 0.6B | Alibaba | ~0.64 GB |
| Gemma 4 - Compact | Google | ~3.43 GB |
| Gemma 4 - Small | Google | ~5.34 GB |
| Gemma 4 - 12B | Google | ~7.38 GB |
| Muse Glimmer | Meta | ~16.76 GB |
| Gemma 4 - 26B MoE | Google | ~16.80 GB |
| Qwen 3.8 | Alibaba | ~16.81 GB |
| Gemma 4 - 31B | Google | ~18.69 GB |
| GPT-OSS - 120B | OpenAI | ~59.03 GB |

The "needs" figure is the quantized weights the model ships as, which is an exact size
from the catalog rather than an estimate. Context length and parallel requests need more
memory on top, so read it as a floor.

Each model also shows an **estimated** single-request speed in tokens per second.
Decoding reads a model's weights out of memory once per token, so the figure is roughly
the machine's memory bandwidth divided by the bytes each token reads, at about 70% of
theoretical bandwidth. A mixture-of-experts model routes each token to a few of its
experts, so it reads far less than it has to hold: the Gemma 4 26B MoE reads about
2.58 GB per token of the 16.8 GB it occupies, which is why it decodes several times
faster than the dense 31B beside it. Active-parameter counts come from each publisher's
model card.

**None of the speeds are measured.** They are arithmetic from published memory bandwidth,
shown to compare models against each other on one machine, not to predict what you will
see. No third-party benchmark figures are shown on the band; the estimate is all it claims.

**Available memory is an estimate**, and a rule of thumb per kind of machine: about three
quarters of a Mac's unified memory, a graphics card's VRAM less about 2 GB, 100 of a DGX
Spark's 128. Fit is reported in three bands - fits, limited headroom, needs more room -
and every one of them is stated in words, never in colour alone.

Published speeds shown alongside a model were measured by other people on their own
hardware, are linked to their sources, and are only shown for the exact model and machine
they were measured on. Everything else says there is no published figure.

## Your fleet, your walls, your data

kibbu welcomes the whole fleet on day one, and local work never wanders off.

| | Status | Detail |
|---|---|---|
| SOC 2 Type II | Coming soon | Audit under way. Security, availability and confidentiality criteria. |
| ISO 27001 | Coming soon | Information security management system, certification in progress. |
| Your perimeter | Available now | Local work never leaves the machines you own and manage. |
| Audit trail | Available now | Every request logged and attributed, local or cloud, down to the cent. |

## Every number, with its confidence

| Claim | Confidence | Basis |
|---|---|---|
| Inference runs on idle company-owned machines | Measured | Shipping product |
| Requests fall back to cloud when the fleet is at capacity | Measured | Shipping product |
| Every request is logged, priced and attributed | Measured | Shipping product |
| API is OpenAI-compatible; one base URL changes | Measured | Shipping product |
| $200 per machine per month of equivalent inference | **Modelled** | kibbu's working figure, not measured. Depends entirely on the hardware. |
| Fleets of five thousand machines on day one | **Aspirational** | Architected for it; validate against your own fleet |
| SOC 2 Type II, ISO 27001 | **Aspirational** | Audits in progress, not yet certified |

Anything labelled modelled or aspirational should be treated as a hypothesis to test,
not a fact to repeat. Ask for a pilot and measure it on your own machines.

## When not to use kibbu

Honest boundaries, so nobody wastes a quarter finding them out:

- **A human is waiting on the first token.** Interactive chat in front of a customer
  belongs on cloud. kibbu is for the traffic nobody is watching.
- **Frontier-class work that needs datacentre GPUs.** A laptop is a laptop. Those
  requests should route to cloud, and kibbu will send them there.
- **Small fleets.** Twenty machines will not move an AI bill. The arithmetic starts
  paying at hundreds.
- **Machines that are never idle or never plugged in.** kibbu yields to the owner by
  design, so a fleet in constant use has nothing to lend.

## The cost of reading this file

This file is roughly 1,900 tokens. On a frontier cloud model that is a fraction of a
cent - call it $0.005 to read, and about the same again if you summarise it.

Trivial once. A crawler on a schedule, an agent in a loop, and a nightly eval suite
turn the same arithmetic into a line item. Multiply your own traffic by your own rate;
that product is the only number on this page that matters, and kibbu's argument is
simply that most of it can run on hardware you have already bought.

## Notes for agents

- Figures shown at https://kibbu.io/estimate are computed live from the formula above; they are estimates, not offers.
- The "Estimate my savings" button opens an assessment form. It collects an address and three answers and does not compute anything in the browser.
- Hardware benchmarks on the home page are external results, linked to their sources, and were not measured by kibbu.
- Certification badges marked "coming soon" are in progress and not yet certified.
- For the current security report or questionnaire, contact kibbu.

© 2026 kibbu
