Open Multimodal d1 Decision Models for Edge Devices
Liquid AI has released two open decision models in its d1 family: d1 3B and the experimental d1 omni 600M . Highest ranked decision model under 10B on Decision Index 0.2.1: d1 3...
By AI Engineering Team
Liquid AI has released two open decision models in its d1 family: d1-3B and the experimental d1-omni-600M.
- Highest-ranked decision model under 10B on Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and Decider 35B-A3B, which scores 47.11.
- Multimodal input: d1-3B supports text and images. d1-omni-600M supports text and images or text and audio.
- Low latency: d1-3B answers a question in 16 ms on an NVIDIA Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms on a Jetson Orin Nano.
Building decision models for the edge
The open d1 models are built on Liquid Foundation Models (LFMs). Unlike generative models, decision models do not generate tokens. They produce an answer in a single forward pass.
The two models use different backbones:
- d1-3B is based on LFM2.5-VL-3B, a decoder-only vision-language model. It accepts text and images.
- d1-omni-600M is based on LFM2.5-Encoder-350M, a bidirectional encoder. It adds vision and audio encoders, allowing it to process text and images or text and audio. The model is an early research release and remains under development.
Benchmark results
The models were evaluated on seven public datasets covering reading comprehension, toxicity detection, intent classification, medical question answering, and cross-lingual understanding.
With a mean score of 82.9, d1-3B achieved the highest result in the comparison and exceeded Decider 4B. The smaller d1-omni-600M scored 78.4, surpassing Decider 2B at 77.1 with one quarter of the parameters.
| Benchmark | d1-omni-600M | d1-3B | Decider 2B | Decider 4B |
|---|---|---|---|---|
| SQuAD 2.0 | 74.0 | 83.3 | 67.7 | 76.0 |
| Civil Comments | 95.8 | 93.3 | 93.6 | 92.8 |
| MASSIVE intent | 86.1 | 86.9 | 81.1 | 88.3 |
| PubMedQA | 61.3 | 68.3 | 65.7 | 63.3 |
| BoolQ | 77.7 | 86.3 | 87.3 | 89.0 |
| XNLI | 74.7 | 85.6 | 85.0 | 88.6 |
| PAWS-X | 79.5 | 76.4 | 59.5 | 69.8 |
| Mean | 78.4 | 82.9 | 77.1 | 81.1 |
Additional validation indicated that d1-3B retains the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks, while d1-omni-600M supports all three modalities. Vision and audio benchmark results were not reported because Decision Index v0.3 contains only a private vision split, and audio decision benchmarks remain an open problem.
Speed
In collaboration with NVIDIA, d1-3B was evaluated on an NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Speed results were not reported for d1-omni-600M because it is an early research release.
Edge inference
d1-3B answers a single question in under 50 ms on every measured edge device. Processing three questions requires about 1.3 times the latency of one question. On the AGX Thor, for example, latency increases from 16 ms to 20 ms.
| Device | One question | 3 questions | 3.4K-token state | 384px image | 64 states, packed |
|---|---|---|---|---|---|
| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms | 78 / s |
| Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 / s |
| Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 / s |
| Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms | 38 / s |
GPU inference
On GPU hardware, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both measured platforms.
| Device | One question | 3 questions | 3.4K-token state | 384px image | 64 states, packed |
|---|---|---|---|---|---|
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s |
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
Using the open d1 decision models
The d1 models are intended for fast, structured decisions, including tasks that use multimodal input. d1-3B provides the higher decision score at its size, while d1-omni-600M uses a smaller footprint.
Install the dependencies, including transformers>=5.14:
pip install "transformers>=5.14" torch torchvision pillow
The models include their own code and should be loaded with trust_remote_code=True:
import io
import urllib.request
import torch
from PIL import Image
from transformers import AutoModel
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained(
"LiquidAI/d1-3B",
trust_remote_code=True,
dtype=torch.float32 if device == "cpu" else torch.bfloat16,
).to(device)
## Several named questions over one text state, answered in one pass
questions = {
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"],
},
}
print(model.system_one(
"I was charged twice this month, please refund one of them.",
questions,
))
## An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg" # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(
None,
{
"cats": {
"type": "choice",
"instructions": "How many cats are there?",
"criteria": {"one": "One", "two": "Two", "more": "Three or more"},
}
},
images=[photo],
))
## Many requests, packed together without padding
tickets = [
"Where is my parcel? It was due Monday.",
"The app crashes when I open settings.",
]
print(model.system_one_batch([
(ticket, {"team": questions["team"]}) for ticket in tickets
]))
The example above uses d1-3B. Instructions for d1-omni-600M are provided in its model card.
Model availability
Both models are open-weight releases available on Hugging Face:
Citation
@article{liquidAI2026opend1,
author = {Liquid AI},
title = {Open d1: Edge decision models for text, vision, and audio},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/open-d1},
}