A Liquid AI bejelentette a d1 döntési modellcsalád első két nyílt forráskódú tagját, a d1-3B-t és a kísérleti d1-omni-600M-et. Ezek a modellek a hagyományos generatív társaikkal ellentétben nem szöveget gyártanak, hanem egyetlen lépésben, azonnali és strukturált döntéseket hoznak meg.
A d1-3B modell képes szöveges és vizuális adatok feldolgozására, és a mérések szerint helyi eszközökön is kevesebb mint 50 milliszekundom alatt válaszol. A kisebb, d1-omni-600M változat ráadásul a hangalapú bemeneteket is kezeli, miközben méretéhez képest kiemelkedő pontosságot nyújt.
Mindkét modell nyílt súlyokkal, szabadon elérhető a Hugging Face felületén keresztül. Használatukhoz a Transformers könyvtár legújabb verziójára van szükség.
Az eredeti szöveg (Hugging Face)
How we built decision models for the edge Benchmark results Speed How to use open d1 decision models Get Started with open d1 decision models Citation Today, we release two open decision models in our d1 decision model family: d1-3B and d1-omni-600M (experimental).
These open d1 decision models are built on our Liquid Foundation Models (LFMs). Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass.
d1-3B and d1-omni-600M are trained from two very different backbones:
We benchmarked d1-3B and d1-omni-600M on seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. d1-3B achieves a mean score of 82.9, the highest in the table and above Decider 4B. d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters.
We validated that d1-3B retains the vision capabilities of its LFM2.5-VL-3B backbone on standard vision benchmarks, and that d1-omni-600M handles all three modalities. We do not report any vision or audio benchmarks, as the Decision Index v0.3 includes only a private vision split and audio decision benchmarks are currently an open problem.
In collaboration with NVIDIA, we evaluated d1-3B on the NVIDIA stack across NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. Since d1-omni-600M is an early research release, we don’t report any speed numbers for it in this release.
Edge inference. d1-3B answers a single question in under 50 ms on every measured device. Three questions take only 1.3x the time of one, with the AGX Thor going from 16 ms to 20 ms.
GPU inference. On GPU, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both platforms.
Reach for d1 decision models when you need fast, structured decisions, including multimodal inputs. d1-3B delivers the highest decision quality at its size, while d1-omni-600M fits where footprint matters.
Install the dependencies (requires transformers>=5.14):
These model ship their own code, so load it with trust_remote_code=True:
For brevity, we only include the example for d1-3B. See the d1-omni-600M model card for instructions on how to run it.
Both decision models are open-weight and available on Hugging Face today:
If you use this work, please cite the release blog: