If you are considering having an assistant trained on your own material, the first honest question is not what it will be able to do. It is whether your material is in a state that can be trained on at all.
Almost nobody knows the answer about their own archive. That is not carelessness. Archives accumulate; nobody sits down and designs one. So this is a plain account of what works, what does not, and what sits awkwardly in between.
What works well
Text that was born as text. Word documents, emails, spreadsheets, notes, exported messages, anything typed. This is the ideal case: the words are already words, the dates are already dates, and nothing has to be guessed.
Documents with consistent structure. Invoices from the same supplier, statements from the same bank, reports in the same template. Repetition is a gift — once the shape is understood, every document of that shape is understood.
Correspondence with a thread. Email chains and message histories carry something single documents do not: sequence. Who said what, when, and what followed. A great deal of what you actually want to ask later is a question about sequence.
Anything with a date on it. Dated material can be put in order, and material in order can answer questions about change — what something used to cost, when a problem first appeared, how often it has recurred.
What does not work
Photographs of paper, taken badly. A phone snap of a document at an angle, in poor light, is often legible to you and not to anything else. Scans are fine. Photographs of a page lying flat are usually fine. Photographs taken at speed generally are not.
Handwriting, mostly. Neat, consistent handwriting can sometimes be read. Ordinary handwriting — annotations in margins, notes to self, a signature over a printed line — usually cannot be read reliably, and unreliable reading is worse than no reading, because it produces confident wrong answers.
Material that contradicts itself. Three versions of the same contract with no indication which is current. A figure that appears as two different numbers in two places. Training does not resolve contradictions; it learns them, and then repeats them back to you with equal confidence.
Anything you are not entitled to use. Material belonging to a client, a former employer, or a third party who has not agreed. This is not a technical limit. It is the first question worth asking, and the answer is sometimes no.
The awkward middle
Most real archives are mostly middle, and this is where the work is.
Duplicates. The same document in four folders under three names. Harmless individually, corrosive in bulk: an assistant that has read something four times can treat it as four separate facts.
Near-duplicates. Worse than duplicates. Draft three and draft four of the same agreement, differing in one clause. Deciding which is authoritative is a judgement about your affairs, not a technical operation, and it cannot be automated away.
Undated material. Usable, but it cannot participate in any question about time, which is a larger loss than it sounds.
Enormous, rarely relevant collections. Fifteen years of newsletters. Every photograph you have taken. Including them is not harmful, but it is expensive in preparation for very little return.
Why this is the first paid step, rather than a promise
It would be easy to say yes to every archive and discover the problems afterwards. We would rather find out first, in writing, and so would you.
That is what the Archive Assessment is. You choose a slice of your material — whichever slice you would most like handled — and we catalogue what is actually in it: what can be trained on, what cannot, what it would take, and what an assistant would be able to do with it once it has. You receive that as a written scope, alongside a working demonstration built on a sample of your own material rather than somebody else's.
You keep both, whether or not you go any further. It is $1,950, it is delivered work rather than the Archive Assessment, and if we can see we cannot serve you well we decline the commission and return it rather than take it.
What to do before you ask anyone
If you want to improve your odds without spending anything:
- Find the largest single folder of typed documents you control outright. That is usually the best starting slice.
- Decide, for any document that exists in several versions, which one is authoritative. Nobody else can decide this.
- Separate anything belonging to someone else. Deal with that question on its own terms.
- Do not reorganise everything first. It is a large job, and the scope will tell you which parts were worth doing.
The last point matters most. The instinct is to tidy the whole archive before anyone looks at it. That instinct costs people months, and most of the tidying turns out to be unnecessary.
Get a written scope of your own material →
Local AI. Private data. Local training.
$8,995 the first year — everything included. $4,995 each year after. The first year costs more because it contains the Jetson Orin placed and configured, your archive loaded and the first training run; every year after is the service running.
Or add the Archive Assessment first — $1,950, credited in full
You order and pay at checkout; we write back within days — a decline returns every dollar, and until our letter confirms the year you may withdraw in writing. From that letter the year is final, and it arrives on the date the letter names. Prefer to write to us first?