What the Jetson Orin in your building can run

Spec sheets are full of numbers that do not answer the question you actually have: what will this thing do for me, in my living room, on my own files, without my work leaving the house? This is an honest map of that. What Digital Twin Pro — the small box your assistant lives in — handles well, what it handles acceptably, and what genuinely belongs on a bigger box or out in the cloud. The short version is that it is very good at the everyday work, and it is not pretending to be a data centre.

For technical readers. The hardware underneath is NVIDIA's Jetson Orin Nano platform: NVIDIA states up to 67 sparse INT8 TOPS (33 dense, Super mode), a 1024-core Ampere GPU, and 8GB of LPDDR5. That is the class of hardware in the original edition; the standard Digital Twin Pro now ships on the Jetson Orin NX 16GB, with up to 157 sparse INT8 TOPS. The rest of this article is what the 67-TOPS class means in practice.

First, what up to 67 sparse INT8 TOPS and 8GB really mean

Two numbers set the boundaries. The up to 67 sparse INT8 TOPS figure describes raw AI compute throughput, which drives how fast tokens come out. The 8GB of unified memory is often the harder limit: the model weights, the key/value cache for context, and the operating system all share that pool. On a device like this, you almost always run quantized models (typically 4-bit) so the weights fit with headroom to spare. That single design choice is what makes a compact, roughly $2-a-month-in-electricity appliance able to run real language models locally by default, so your data stays on the device.

Large language models: the 7-8B class is the sweet spot

Quantized 7-8B-parameter models are exactly what this hardware was tuned for. Think of small open models at reduced precision. They fit comfortably in 8GB and leave room for a usable context window.

Realistic, approximate throughput ranges for a single conversation on this class of hardware:

See what your own private AI can do

The first year is the whole price, and every year after is the service running. The machine stays ours to place, maintain and replace; the model trained on your material is yours alone while the service runs. See what it costs →

  • 3-4B models (4-bit): roughly 20-40 tokens/sec (approximate) — fast enough to feel snappy for chat and drafting.
  • 7-8B models (4-bit): roughly 8-18 tokens/sec (approximate) — comfortably faster than most people read, good for assistants, summarization, Q&A over your notes, and agent tasks.
  • Time-to-first-token grows with how much context you feed in; short prompts feel near-instant, very long documents take longer to "warm up."

These are ballpark figures; real speed varies with model, quantization, prompt length, and workload. Your assistant ships preinstalled inside the preconfigured software environment and is designed to drive exactly these local models — including the Nemotron family — so you get a working assistant out of the box rather than a bare runtime.

What about bigger models?

Models in the 13-14B range can sometimes be squeezed in at aggressive quantization, but you trade away context room and speed, and quality can suffer from the heavier compression. It is doable for experimentation, not something we would call a great daily-driver experience. Beyond that, 30B+ and frontier-scale models simply do not fit in 8GB — that is a cloud or bigger-tier job, and we are upfront about it.

Speech: ASR and TTS run well locally

Speech is one of the strongest use cases for this appliance. Small and medium automatic speech recognition (ASR) models — the kind used for transcription and voice commands — run efficiently, often comfortably faster than real time for the smaller variants, meaning a minute of audio transcribes in well under a minute. Text-to-speech (TTS) for a responsive local voice assistant is also well within reach.

Because ASR, a 7-8B LLM, and TTS can all live on the device, you can build a local voice loop: speak, transcribe, reason, and reply — and nothing leaves your home unless you choose to send it. That is a genuinely private smart-assistant pattern that is hard to get from cloud-tethered speakers.

Vision: small models, yes; giant multimodal, no

Small computer-vision models are a good fit — image classification, object detection, and similar tasks run well, which is unsurprising given the Jetson platform's roots in edge vision and robotics. Compact vision-language models can also run for basic image description and visual Q&A.

The honest caveat: large, state-of-the-art multimodal models and heavy real-time video analytics push past what 8GB and up to 67 sparse INT8 TOPS handle gracefully. For high-frame-rate, many-stream, or frontier-quality vision work, you want more memory and more compute than this tier provides.

Where the cloud still wins — and how your own account helps

We would rather set expectations correctly than oversell. The cloud still leads clearly in a few areas:

  • Frontier-scale reasoning: the largest, most capable models do not fit locally and outperform 7-8B models on the hardest tasks.
  • Very large context windows: stuffing hundreds of thousands of tokens into a single prompt is memory-hungry and better suited to cloud infrastructure.
  • Peak throughput and heavy concurrency: serving many simultaneous users at high speed is a data-center job.

Digital Twin Pro deals with this simply: you can connect an outside model you already pay for and it's optional. Your everyday, private, high-volume work stays on the shelf and costs you nothing extra. When you deliberately want one of the big outside services, you send that one job out under your own account. You decide, job by job, and the box shows you when a job went out.

The bottom line

A box this size is not trying to be a data centre, and it does not need to be. For the work most people actually want running privately and all the time — drafting, summarising, chat, listening and speaking, answering from your own papers, simple picture work — it is capable, sips power, arrives fully assembled, and does all of it on your own shelf for a few dollars a month in electricity. The first year is the whole price, and every year after is the service running. The machine stays ours to place, maintain and replace; the model trained on your material is yours alone while the service runs.

Frequently asked questions

What size LLM can the Jetson Orin Nano run?

Quantized 7-8B-parameter models (typically 4-bit) are the sweet spot and fit comfortably in the 8GB of memory. Smaller 3-4B models run faster, 13-14B is possible at aggressive quantization with trade-offs, and 30B+ or frontier-scale models do not fit and are better run in the cloud.

How many tokens per second should I expect?

These are approximate ranges for a single conversation: roughly 20-40 tokens/sec (approximate) for 3-4B models and roughly 8-18 tokens/sec (approximate) for 7-8B models, both at 4-bit. Actual speed varies with the specific model, quantization, and prompt length, so treat these as ballpark figures rather than benchmarks.

Can it run speech and vision, not just chat?

Yes. Small and medium ASR (transcription) and TTS models run well locally, and small vision models like image classification and object detection are a good fit. Large state-of-the-art multimodal models and heavy real-time video analytics exceed what this tier handles gracefully.

What genuinely needs the cloud?

Frontier-scale reasoning, very large context windows, and high-concurrency peak throughput still favor the cloud. Digital Twin Pro supports bring-your-own-key (your own account) so you can optionally call cloud models with your own key when you choose, while keeping everyday work local and private by default.

Do local models keep working if I cancel the service?

The first year is the whole price, and every year after is the service running. The machine stays ours to place, maintain and replace; the model trained on your material is yours alone while the service runs.

Local AI. Private data. Local training.

$8,995 the first year — everything included. $4,995 each year after. The first year costs more because it contains the Jetson Orin placed and configured, your archive loaded and the first training run; every year after is the service running.

Begin the first year — $8,995

Or add the Archive Assessment first — $1,950, credited in full

You order and pay at checkout; we write back within days — a decline returns every dollar, and until our letter confirms the year you may withdraw in writing. From that letter the year is final, and it arrives on the date the letter names. Prefer to write to us first?

One decision, and then it is handled

One price, and one way to begin.


The same assistant at either address, for the same price. Begin with the first year; if you like, add the Archive Assessment before it — delivered work you keep, credited in full.

The first year

$8,995

Everything included. $4,995 each year after.

Begin the first year — $8,995

Or add the Archive Assessment first — $1,950, credited in full

The first year costs more because it contains the Jetson Orin placed and configured, your archive loaded and the first training run; every year after is the service running. You order and pay at checkout; we write back within days — a decline returns every dollar, and until our letter confirms the year you may withdraw in writing. From that letter the year is final, and it arrives on the date the letter names.

We answer in writing, within one business day — and if we cannot serve you well we say so before our letter confirms a year. Prefer to write to us first?