Technology

How Hatek Lingua AI processes a video.

The product combines speech recognition, translation, generated speech, video processing and quality evaluation in an end-to-end workflow. Each part contributes to the final translated video.

This page describes the product architecture used to connect speech, translation, generated voice, video processing and quality review in one workflow.

Processing pipeline

The system starts with the original video and ends with a translated audiovisual version.

The main technical challenge is keeping language meaning, speech timing and visible mouth movement consistent with one another.

01

Speech recognition and timing

The system identifies spoken words, speaker changes and the timing of each segment. This timing information is needed later when translated speech is generated.

02

Context-aware translation

The spoken content is translated into the target language while preserving meaning, names, terminology and conversational context as closely as possible.

03

Target-language speech generation

The translated text is converted into speech. Pronunciation, naturalness, pacing and duration are important because the audio must still fit the video.

04

Visual speech synchronisation

The video is adjusted so the visible speaker's mouth movement follows the generated target-language speech while the surrounding scene remains stable.

05

Quality evaluation

The output is checked for translation quality, pronunciation, speech timing, visual consistency and common failure cases before it is accepted.

06

One workflow for the user

Hatek Lingua AI hides the complexity of the individual AI systems behind a simple workflow: upload video, choose an available target language, review the result and export.

Product workflow

The processing pipeline is reflected directly in the product interface.

Users move from source media and language settings to generated output, review and export without leaving the localisation workspace.

Hatek Lingua AI video translation workspace
Video translation

Source video, language, voice and output controls in one view.

The video workspace brings together target-language selection, generated speech, subtitles, speaker handling, lip synchronisation and export settings.

Hatek Lingua AI dubbing and lip-sync workspace
Dubbing & lip-sync

Review original and translated output side by side.

The dubbing workspace is designed for comparing the source and dubbed versions while adjusting language, voice and synchronisation settings.

The product is offered through managed access for customer, partner and language workflows.

What quality means

A translated video is only useful when several kinds of quality are acceptable at the same time.

A strong translation with poor pronunciation is not enough. Good audio with unstable lip synchronisation is not enough either. Evaluation therefore has to cover the full audiovisual result.

01Meaning

Does the translated speech preserve the intended message?

02Pronunciation

Are names, local words and target-language sounds spoken correctly?

03Timing

Does the translated speech fit the available speaking time naturally?

04Visual consistency

Does the face remain stable while mouth movement follows the new speech?

Workspace controls

Product behaviour can be configured around teams and localisation workflows.

Beyond the translation pipeline, the product includes account, language, notification, storage and translation preferences.

These controls help organisations standardise how projects are created and managed while keeping account and workspace settings in one place.

Hatek Lingua AI settings and workspace preferences
Settings

Account, translation, language and workspace preferences.

The settings area centralises defaults for translation, regional preferences, notifications, storage and account controls.

AI infrastructure

GPU acceleration supports compute-intensive speech, language and video workloads.

Hatek Lingua AI is designed to use NVIDIA GPU infrastructure for compute-intensive speech, language and video workloads.

NVIDIA CUDA can accelerate model workloads, while TensorRT can optimise inference for latency and cost as the platform scales. Infrastructure choices are benchmarked against model size, throughput, quality and operating cost.

Technical collaboration

We work with compute, data, language and technical partners.

Partnership opportunities include language data, native-language expertise, AI infrastructure, video workflows and technical integration.

Discuss collaboration