The inference runtime
Configuration
Every environment variable, and the food-database (FOOD_SOURCE) options
Everything is environment variables, validated at boot; a bad value stops the process rather than degrading silently. The annotated master list is .env.example. The openplate docs list every variable for this container on one page. They appear next to those for the app and the sync service. The most important variables are:
| variable | default | ||||
|---|---|---|---|---|---|
MODEL_PROFILE | lite | lite \ | lite-apache \ | quality \ | external |
API_KEYS | (generated) | Comma-separated bearer keys. Set this. | |||
PORT | 8300 | The only published port. | |||
CONCURRENCY | 2 | In-flight scans; also sets llama.cpp's KV slots. It does not add CPU threads: the slots share the LLAMA_THREADS threads. | |||
MAX_QUEUE_DEPTH | 8 | Past this, callers get 429 + Retry-After. | |||
RATE_LIMIT_RPM | 60 | Per key. | |||
LATENCY_CEILING_MS | 0 | 0 = disabled. Admission policy: refuse work you cannot finish in time. See Hardware. | |||
RUNTIME_COMPLETION_TIMEOUT_MS | 600000 | Total bound on one completion call; 0 = disabled. Liveness, not latency policy: it releases a worker slot a wedged runtime will never return. Applies in bundled mode too. See Runtimes. | |||
IMAGE_MAX_LONG_EDGE | 896 | Downscale target. Latency rises with the square. | |||
MAX_IMAGE_BYTES | 8388608 | The largest photo size accepted after base64 decoding, in bytes (8 MiB). A larger file is rejected with a message to resize it. The request body limit is derived from this value. | |||
FOOD_SOURCE | fdc | See Food data. | |||
CONTEXT_SIZE | 8192 | Context per in-flight scan. The container multiplies it by CONCURRENCY before handing it to llama.cpp, because llama.cpp's -c is the total it splits across slots. | |||
LLAMA_EXTRA_ARGS | (empty) | Extra flags added to the end of the llama-server command, split at spaces. Bundled mode only. | |||
LLAMA_THREADS | nproc - 2 | CPU threads for llama.cpp (its -t). Two cores are left for the service, image decode, and the OS, so a 6-core box runs 4 threads and a 4-core box 2, whatever CONCURRENCY says. Giving llama.cpp every core makes the box contended, not faster. The startup log prints the value as -t N. Bundled mode only. | |||
MODELS_DIR | /models | The weights volume. | |||
RUNTIME_PORT | 8080 | The port of the bundled llama-server, on 127.0.0.1 inside the container. Bundled mode only. | |||
WEIGHTS_MIRROR_BASE | (empty) | Optional mirror; Hugging Face is the fallback. | |||
GPU_LAYERS | (auto) | Override the GPU auto-detect. 0 forces CPU. | |||
NVIDIA_VISIBLE_DEVICES | (set by the runtime) | The NVIDIA container runtime sets it when you pass --gpus all. Auto-detect reads it. Any value other than void or none offloads every layer. You do not set it yourself. | |||
LOG_LEVEL | info | debug \ | info \ | warn \ | error |
PROFILE | (from MODEL_PROFILE) | The profile name in the start-up log: lite, quality or custom. The container sets it from MODEL_PROFILE. Setting it yourself changes only that log line. |
External-mode variables (MODEL_RUNTIME_URL, MODEL_ID, MODEL_RUNTIME_API_KEY) are documented in Bring your own runtime.
llama-server runs inside the container bound to 127.0.0.1 only and is not reachable from outside it. That is not configurable: it is an unauthenticated raw vision endpoint, and the point is that you cannot publish it by accident.
Food data (FoodSource)
The model identifies foods and estimates grams. Macros are resolved from a food database, by name, never generated by the model.
FOOD_SOURCE | what it does | network | notes |
|---|---|---|---|
fdc (default) | Looks names up in a bundled extract of USDA FoodData Central, 8,041 generic foods, shipped inside the image at data/fdc-foods.json. | none | Offline, no key, no account, no outbound request. Public domain. It is the default because it is the only option that needs nothing from anybody. |
off | Queries Open Food Facts live at your runtime. | outbound, per scan | Strong on branded and packaged products, weaker on generic cooked food. The address is OFF_API_URL, https://world.openfoodfacts.org by default. Read the licence note below before enabling. Nothing OFF-derived ships in this image. |
lcc | Queries the public lowcarbcheck API. | outbound, per scan | The broadest data of the three (curated + BLS + USDA), and remote-only permanently, because BLS 4.0 forbids redistribution. Attribution is passed through to the response so it reaches the UI. Without LCC_API_KEY, every request runs on LowCarbCheck's free anonymous tier; see below. |
none | No resolution. Every item comes back with null macros. | none | For clients that do their own nutrition lookup. |
-e FOOD_SOURCE=fdc # default
-e FDC_DATASET_PATH=./data/fdc-foods.json # relative to the working directory
-e OFF_API_URL=https://world.openfoodfacts.org # only read when FOOD_SOURCE=off
-e LCC_API_URL=https://lowcarbcheck.org # only read when FOOD_SOURCE=lcc
-e LCC_API_KEY=lcc_live_… # optional; only read when FOOD_SOURCE=lcc
-e EMBEDDING_RUNTIME_URL=http://… # optional; enables hybrid re-rankingWithout LCC_API_KEY, FOOD_SOURCE=lcc runs on LowCarbCheck's anonymous tier: 1,000 credits per UTC day, shared by all requests from your IP address. A search costs 1 credit. A scan issues up to 3 search queries per identified item, set by the refinement cap in search-foods.ts, across up to 8 items. This makes the worst case 24 credits per scan. The pipeline never calls the per-food endpoint, so that covers the entire cost. That worst case permits about 41 scans a day. Most days permit more, because a search stops as soon as one query clears the accept threshold. When the tier runs out, every remaining item resolves to null macros until the next UTC day, and the scan still returns 200. A free key from lowcarbcheck.org/developers raises the allowance to 100,000 credits a month at up to 120 requests a minute. This allowance belongs to the key rather than your IP address. Set it as LCC_API_KEY. The service sends it as a bearer token to LCC_API_URL and nowhere else, and never logs it. A key LowCarbCheck rejects fails the same way an empty allowance does: null macros and a 200. fdc needs no network and no allowance.
Resolved macros are labelled. Foods matched against the database include a provenance of "corpus", plus an attribution string when required by the source. Unmatched foods omit both fields, and macrosPer100g is null. openplate exposes these values, so users can separate confirmed database entries from items with no macro data.
A missing food database does not stop scans. If FDC_DATASET_PATH points at nothing, the service logs a warning, disables resolution, and keeps identifying plates. You get names and grams with null macros: degraded, not broken. Regenerate the extract with pnpm food-data:fdc (needs network).
EMBEDDING_RUNTIME_URL is optional. Point it at a second OpenAI-compatible runtime serving /v1/embeddings (e.g. llama-server --embedding) and retrieval becomes hybrid: the lexical scorer finds candidates and the embedding model re-ranks them, so "grilled chicken thigh" lands on the right row even when the database words it differently. Leave it unset, the default, and retrieval is lexical-only, which is a slightly worse ranking and never an error. An unreachable embedding runtime degrades to lexical-only with one warning; it never fails a scan.
FOOD_SOURCE=off and ODbL share-alike
Open Food Facts data is licensed under the Open Database Licence (ODbL), which is share-alike. If you enable this connector and then publish or redistribute a database that incorporates OFF data (not the individual lookups you display, but a derived database), the ODbL obliges you to make that derived database available under the ODbL as well, and to attribute Open Food Facts.
For a household instance that displays a lookup and stores it in your own diary, this does not apply. If you are building a product on top of this, it does, and that is why off is not the default. fdc carries no share-alike obligation.
Model and data licence terms are collected in Licensing.