The landscape of on-device machine translation is undergoing a fundamental architectural shift. Historically, integrating local translation capabilities into mobile and edge software required a sprawling collection of separate, dedicated models for every individual language pair. Developers seeking to support English-to-French, English-to-German, and countless other combinations were forced to bundle numerous distinct binary files into their applications. At a global scale—spanning thousands of potential language directions—this traditional approach quickly becomes completely unsustainable, bogging down device storage, draining battery life, and creating massive deployment hurdles for software engineers.
Faced with this bottleneck, developers have traditionally been forced to make a difficult compromise. They could either route sensitive translation requests through remote cloud servers to ensure fast and accurate results, or they could restrict their applications to local processing with severely limited language support. Now, however, the AI Research team at Tether has introduced a breakthrough family of multilingual translation models known as TranslatePsy-EuroNano. This new offering completely reimagines edge-based translation by leveraging a pair of comprehensive multilingual models instead of relying on an unwieldy web of separate bilingual binaries.
This innovative approach makes it feasible to support a full European market directly on-device, bypassing the traditional requirement of bundling dozens of separate model files that render mobile apps bloated and impractical. By setting a new benchmark for computational efficiency, translation quality, and processing speed, Tether AI’s open-source edge translation models are opening up entirely new possibilities for software developers across industries.
What Makes This Possible
At the core of Tether’s technological achievement is a clever architectural design that utilizes English as a central pivot language. Rather than training distinct models for every permutation, the research team trained models from scratch using independently curated and preprocessed open-source parallel data, entirely avoiding the use of pretrained model checkpoints for initialization.
Using English as a pivot, just two multilingual models enable seamless translation across a total of ten languages—comprising English and nine major European languages. This setup successfully generates ninety possible translation directions, encompassing English-to-European, European-to-English, and direct European-to-European translations.
When measured against existing open-source standards, the efficiency gains are staggering. Tether’s deployment is remarkably compact, taking up between 36 megabytes and 89 megabytes of storage depending on the selected performance tier. By comparison, an equivalent setup utilizing Mozilla Firefox’s Bergamot-based translation system requires eighteen separate bilingual models that total a bulky 633 megabytes to provide the exact same language coverage. At its smallest tier, Tether’s deployment achieves an astonishing footprint that is up to 17.6 times smaller than previous on-device approaches while maintaining a strikingly comparable level of translation quality.
These models are small enough to run efficiently on resource-constrained edge devices while simultaneously supporting nine European languages from a single multilingual deployment. This makes robust multilingual experiences practical for a vastly wider range of software applications. Potential use cases span multiple sectors, including travel and navigation apps, offline educational platforms that present localized lessons and resources directly on-device, and specialized tools designed for academics and field researchers.
Furthermore, because the model weights are made openly available, researchers and enterprise developers can fine-tune them for specialized domains. These customized iterations can then be applied to customer support chatbots, offline virtual assistants, educational AI tutors, and sophisticated question-answering systems, before being published for the broader developer community to build upon. Rather than functioning as a single, isolated translation utility, these models are engineered to serve as a foundational bedrock for an entire ecosystem of multilingual AI systems.
No Third-Party Involvement
The implications of shifting heavy processing workloads entirely to the edge extend far beyond mere storage optimization and performance metrics. To understand the broader value of local translation, one only needs to consider the typical user experience of traveling abroad and utilizing a standard mobile translation app.
When a user types a phrase into a traditional translation app and watches the translated text materialize seconds later, they rarely consider the complex journey their data undertakes during that brief half-second interval. Before a translation appears on a smartphone screen, sensitive text typically leaves the device in the form of an encrypted data packet, gets routed to a regional network server, and is ultimately processed within a centralized corporate data center. The only exception to this data-harvesting pipeline has traditionally been offline translation modes, which often suffer from severe limitations in accuracy and vocabulary.
When translation processing runs entirely on-device, however, user text never leaves the physical hardware of the user’s device. There are no intermediary third-party servers involved in the transaction, no complex or intrusive data-processing agreements required, and no cross-border data transfers that could trigger compliance headaches. This local-first paradigm ensures absolute privacy and provides significantly greater clarity and control over how user data is handled.
How Tether’s Models Work
A closer examination of the underlying methodology reveals how Tether achieves such dramatic compression without sacrificing linguistic fidelity. Traditional translation architectures rely heavily on dedicated bilingual models. For example, Firefox’s Bergamot-based translator requires a completely separate model for every single language direction. Covering nine European languages in both directions necessitates loading eighteen distinct models, resulting in the aforementioned 633 megabyte disk footprint.
In sharp contrast, Tether’s approach utilizes just two multilingual checkpoints per tier—one for each primary translation direction—covering all nine languages with a single, efficient load operation. Loading fewer and significantly smaller models inherently translates to faster execution times and quicker system responses.
Controlled CPU benchmarking conducted on the FLORES-200 dataset highlights these performance advantages clearly. When tasked with returning its first translated sentence, the Firefox-based system took 10.8 seconds. Meanwhile, Tether’s lightweight "TinyQ" tier returned the initial sentence in just 4.2 seconds—more than twice as fast—while the higher-performance "BaseQ" tier completed the task in 6.8 seconds.
The BaseQ tier, which represents the top performance offering in the lineup, retains an impressive 98.4 percent of the translation quality achieved by Meta’s massive NLLB-200 model when translating into English, while closely tracking Firefox’s proprietary scores despite utilizing only a fraction of the storage space. A slight performance gap does widen when translating outward from English, a phenomenon that is typical for multilingual models that share a single decoder across multiple languages rather than dedicating a discrete model to each individual language pair.
Other Translation Models
Tether’s European language release is part of a broader, ambitious initiative by Tether AI Research to deliver the most efficient, highest-quality, and fastest open-source multilingual translation models built specifically for the edge. Recognizing the persistent digital divide and the severe underinvestment in artificial intelligence across the African continent—a barrier that has historically hindered technological adoption for over a billion people—the research team has simultaneously released a dedicated set of models for African languages under the banner of TranslatePsy-AfriSLM.
TranslatePsy-AfriSLM is a comprehensive collection of open-source machine translation resources specifically engineered to cover nineteen Sub-Saharan African languages. Crucially, these specialized models outperform much larger, established systems such as Google’s TranslateGemma and Meta’s NLLB.
For years, existing open-source large language models have chronically underperformed when applied to African machine translation tasks. This systemic failure has largely been driven by a severe shortage of large-scale, high-quality, open-source parallel data, which has historically constrained the development of competitive small language models across the continent. By curating targeted datasets and engineering highly compressed architectures, Tether is addressing this critical gap.
Making advanced language models fully accessible on-device while rigorously maintaining high standards of efficiency, quality, and speed has long presented a formidable engineering challenge. Keeping artificial intelligence local and ensuring that foundational technologies remain universally accessible is a core pillar of Tether’s ongoing operational mission.
The newly developed European translation models, alongside the TranslatePsy-AfriSLM collection, are available immediately through the QVAC software development kit. These resources are designed for frictionless integration across a wide array of operating systems and hardware environments, including Android, iOS, Linux, macOS, and Windows. Developers and researchers interested in auditing the source code, examining the technical documentation, or beginning local integrations can access the complete project repositories directly through the official QVAC GitHub portal.