"up to 67 sparse INT8 TOPS (33 dense, Super mode)" looks great on a spec sheet, but the number most buyers actually care about is different: what can this thing run in my living room, on my own data, without a cloud bill? Digital Twin Pro is powered by NVIDIA's Jetson Orin Nano platform (up to 67 sparse INT8 TOPS, a 1024-core Ampere GPU, and 8GB of LPDDR5), and that hardware has a clear sweet spot. Below is an honest, technical map of what runs well, what runs acceptably, and what genuinely belongs on a bigger tier or in the cloud.
First, what up to 67 sparse INT8 TOPS and 8GB really mean
Two numbers set the boundaries. The up to 67 sparse INT8 TOPS figure describes raw AI compute throughput, which drives how fast tokens come out. The 8GB of unified memory is often the harder limit: the model weights, the key/value cache for context, and the operating system all share that pool. On a device like this, you almost always run quantized models (typically 4-bit) so the weights fit with headroom to spare. That single design choice is what makes a compact, roughly $2-a-month-in-electricity appliance able to run real language models locally by default, so your data stays on the device.
Large language models: the 7-8B class is the sweet spot
Quantized 7-8B-parameter models are exactly what this hardware was tuned for. Think of the well-known open models in the 7B, 8B, and small Mistral or Qwen families at 4-bit quantization. They fit comfortably in 8GB and leave room for a usable context window.
Realistic, approximate throughput ranges for a single conversation on this class of hardware:
See what your own private AI can do
Digital Twin Pro runs real assistants entirely at home. One box, $1,699, no subscription.
Explore Digital Twin Pro →- 3-4B models (4-bit): roughly 20-40 tokens/sec (approximate) — fast enough to feel snappy for chat and drafting.
- 7-8B models (4-bit): roughly 8-18 tokens/sec (approximate) — comfortably faster than most people read, good for assistants, summarization, Q&A over your notes, and agent tasks.
- Time-to-first-token grows with how much context you feed in; short prompts feel near-instant, very long documents take longer to "warm up."
These are ballpark figures; real speed varies with model, quantization, prompt length, and workload. Hermes, which ships preinstalled (OpenClaw optional), inside the preconfigured software environment is designed to drive exactly these local models, so you get a working assistant out of the box rather than a bare runtime.
What about bigger models?
Models in the 13-14B range can sometimes be squeezed in at aggressive quantization, but you trade away context room and speed, and quality can suffer from the heavier compression. It is doable for experimentation, not something we would call a great daily-driver experience. Beyond that, 30B+ and frontier-scale models simply do not fit in 8GB — that is a cloud or bigger-tier job, and we are upfront about it.
Speech: ASR and TTS run well locally
Speech is one of the strongest use cases for this appliance. Small and medium automatic speech recognition (ASR) models — the kind used for transcription and voice commands — run efficiently, often comfortably faster than real time for the smaller variants, meaning a minute of audio transcribes in well under a minute. Text-to-speech (TTS) for a responsive local voice assistant is also well within reach.
Because ASR, a 7-8B LLM, and TTS can all live on the device, you can build a local voice loop: speak, transcribe, reason, and reply; data stays unless you enable. That is a genuinely private smart-assistant pattern that is hard to get from cloud-tethered speakers.
Vision: small models, yes; giant multimodal, no
Small computer-vision models are a good fit — image classification, object detection, and similar tasks run well, which is unsurprising given the Jetson platform's roots in edge vision and robotics. Compact vision-language models can also run for basic image description and visual Q&A.
The honest caveat: large, state-of-the-art multimodal models and heavy real-time video analytics push past what 8GB and up to 67 sparse INT8 TOPS handle gracefully. For high-frame-rate, many-stream, or frontier-quality vision work, you want more memory and more compute than this tier provides.
Where the cloud still wins — and how BYOK helps
We would rather set expectations correctly than oversell. The cloud still leads clearly in a few areas:
- Frontier-scale reasoning: the largest, most capable models do not fit locally and outperform 7-8B models on the hardest tasks.
- Very large context windows: stuffing hundreds of thousands of tokens into a single prompt is memory-hungry and better suited to cloud infrastructure.
- Peak throughput and heavy concurrency: serving many simultaneous users at high speed is a data-center job.
Digital Twin Pro handles this pragmatically with bring-your-own-key (BYOK). Your everyday, private, high-frequency work stays local and free of cloud costs; when you deliberately want a frontier model, you can point the device at a cloud API using your own key. You decide, per task, when data leaves the device — the default is local and private.
The bottom line
A up to 67 sparse INT8 TOPS, 8GB appliance is not trying to be a data center, and it does not need to be. For quantized 7-8B LLMs, local speech, and small vision models — the workloads most people actually want running privately and continuously — it is a capable, low-power, fully-assembled option that runs entirely on your own hardware for roughly $2 a month in electricity. There are no required ongoing fees — optional one-time services (setup help, Software Refresh, Life Upload) are available whenever you want them.
Frequently asked questions
What size LLM can the Jetson Orin Nano run?
Quantized 7-8B-parameter models (typically 4-bit) are the sweet spot and fit comfortably in the 8GB of memory. Smaller 3-4B models run faster, 13-14B is possible at aggressive quantization with trade-offs, and 30B+ or frontier-scale models do not fit and are better run in the cloud.
How many tokens per second should I expect?
These are approximate ranges for a single conversation: roughly 20-40 tokens/sec (approximate) for 3-4B models and roughly 8-18 tokens/sec (approximate) for 7-8B models, both at 4-bit. Actual speed varies with the specific model, quantization, and prompt length, so treat these as ballpark figures rather than benchmarks.
Can it run speech and vision, not just chat?
Yes. Small and medium ASR (transcription) and TTS models run well locally, and small vision models like image classification and object detection are a good fit. Large state-of-the-art multimodal models and heavy real-time video analytics exceed what this tier handles gracefully.
What genuinely needs the cloud?
Frontier-scale reasoning, very large context windows, and high-concurrency peak throughput still favor the cloud. Digital Twin Pro supports bring-your-own-key (BYOK) so you can optionally call cloud models with your own key when you choose, while keeping everyday work local and private by default.
Do I need the monthly plan to run models locally?
No. Local models run on the device with no required subscription. After the $1,699 purchase there are no required ongoing fees; any service you ever add is a one-time purchase.
Ready to see it on your own desk? Explore Digital Twin Pro Edge — from $1,699 or compare the systems.