On-premise & data residency
Some content can't leave the building. Fily can move in.
Run the same pipeline inside your own infrastructure — in the US or the EU, on open-weight models — so source files, translations and terminology never cross a network boundary you don't control.
Three levels of isolation
“On-premise” means different things to different security teams. These are the three we actually deploy — pick the one that answers your requirement, not the strictest one on the list.
Option 1
Data residency
Our infrastructure, pinned to a region. Your content is processed and stored in the US or the EU and nowhere else.
- Fastest to set up — no infrastructure work on your side
- Region-pinned processing and storage
- Frontier models, full feature set
- Dedicated worker pool available
Best when the requirement is jurisdictional, not architectural.
Option 2
Private cloud
The pipeline deployed into your own cloud account — your VPC, your storage buckets, your keys. We operate it; you own the perimeter.
- Runs in your AWS / Azure / GCP account
- Content never leaves infrastructure you control
- You choose which model provider it may call, if any
- Your existing network controls, logging and IAM apply
Best when your security team needs the data inside your boundary.
Option 3
Fully on-premise
Fily inside your datacenter, running open-weight models on your own GPUs. No outbound calls for translation at all.
- Open-weight models served locally — no third-party model API
- Open-weight ASR for audio instead of hosted engines
- Works in restricted-egress networks
- You supply the GPU capacity; we size it with you
Best when no content may leave the building, by policy or by law.
What doesn't change
The pipeline is the product, and the pipeline is ours — not a wrapper around someone else's API. That is why it travels.
- The 12-step QA pipeline, step for step
- Dedicated codecs per CAT format — Trados, memoQ, Phrase, Wordfast, XLIFF, RTF
- Glossary enforcement with the 4-layer cascade and tag-safe rollback
- Translation memory pass-through and your Golden TM
- Tag repair and round-trip verification before delivery
- The QA report attached to every delivery
- The browser review editor, quality scoring and reviewer effort reporting
What does change — plainly
Open-weight models you host are not the same as the frontier models our cloud uses. On most business, legal and healthcare content the gap is small and the QA steps absorb much of it — but it is a real gap, and anyone who tells you otherwise is selling.
So we do not ask you to take that on faith. Before you commit to an on-premise deployment, we run your content through both configurations and give you the comparison — same files, same glossary, same QA report, scored side by side. If the difference matters for your use case, you will see it before you sign, not after.
- • Audio uses open-weight ASR instead of the hosted engines
- • GPU capacity is yours to provide; we size it with you
- • Model updates ship on your maintenance window, not ours
How a deployment goes
Scoping call with your security and infrastructure people
Volumes, formats, languages, where the boundary has to be, and which of the three options actually answers your requirement.
A benchmark on your own content
We run a representative sample through the hosted and the open-weight configuration and hand you both outputs with their QA reports.
Sizing and quote
GPU footprint, throughput, and a fixed deployment fee plus a support agreement. No per-seat licensing.
Install, verify, hand over
We deploy, run acceptance tests against your files, and train your team. Support and updates continue under the agreement.
Tell us where the boundary has to be.
On-premise deployments are quoted individually — there is no self-serve option and we would rather scope it properly than sell you a tier.