Skip to content
EN

Capabilities

Linguistic engineering and knowledge-based generation

We design systems capable of analysing, representing, transforming and generating language under explicit constraints of meaning, terminology, evidence, audience and context.

On one page

We combine NLP, computational terminology, semantic representation, contrastive research, localisation, NLG planning, structured generation and generative models. We design systems in which language works as a computational layer connected to data and knowledge, with controls over what may be transformed and what must be preserved.

  1. 01Linguistic analysis and representation. Language does not reach the system as a clean structure.
  2. 02NLG planning and generation. Generating reliable language begins before any writing.
  3. 03Controlled generation. Not every linguistic decision belongs to the model.

We work with language as a computational layer of the system: from extraction and semantic representation through to NLG, multi-variety localisation, terminology, adaptation by audience and automatic evaluation.

We combine linguistic and generative models with knowledge structures, rules, schemas, terminology memories, validators and specialist review.

The aim is not to produce plausible text. It is to produce language that preserves what the system knows, respects what it cannot change and can be evaluated afterwards.

Linguistic analysis and representation

Language does not reach the system as a clean structure. It contains entities, relationships, events, temporal references, quantities, modality, negation, conditions, arguments, uncertainty, terminology and contextual dependencies that may need an explicit representation.

We combine linguistic models, structured extraction, classification, embeddings, retrieval, rules and validators to turn open language into computable objects. When the domain requires it, those representations are connected to taxonomies, terminologies, conceptual schemas and knowledge structures.

We are not only trying to understand a sentence: we are trying to represent its meaning well enough for another component to reason, compare, generate or act upon it.

NLG planning and generation

Generating reliable language begins before any writing. When the problem demands it, we separate the available evidence, the representation of facts and relationships, content selection, narrative planning and linguistic realisation.

On that architecture we combine deterministic computation, rules, templates, retrieval, generative models and structural constraints. The system can decide what to communicate, in what order, at what granularity, for which audience, in which language and through which linguistic form, without necessarily delegating all of those decisions to a single model.

One and the same representation of evidence can feed different narratives, languages, durations and audience profiles while preserving a common factual core.

Generative models hugely expand the space of possible expression; NLG engineering determines how much freedom they should have and where that freedom must end.

Controlled generation
  1. Sources
  2. Evidence
  3. Representation
  4. Planning
  5. Realisation
  6. Validation

Controlled generation

Not every linguistic decision belongs to the model.

Deterministic

Figures · calculations · identifiers · units · mandatory relationships · invariants

Constrained

Terminology · structure · length · audience · language · register · format

Generative

Writing · explanation · reformulation · composition · style

Evaluated

Fidelity · coverage · contradictions · naturalness · terminology · invariants

Computational localisation and linguistic variation

Localising does not consist of replacing the expressions of one variety with those of another. It requires modelling simultaneously what must change, what may change and what must remain unaltered.

We work with lexical, morphosyntactic, terminological, pragmatic and register differences; linguistic memories and constraints; contrastive research; corpus generation; semantic evaluation and adversarial review.

A transformation can be perfectly natural and yet introduce an inference, remove a condition, alter a relationship or shift the sense. That is why we treat localisation as a problem of controlled transformation, with context, evidence, constraints and explicit evaluation criteria.

Quality does not consist of the result «sounding local». It consists of changing exactly what must change without modifying what must remain true.

Computational terminology and conceptual knowledge

A word is not a concept, and a string match does not prove identity of meaning.

We build terminological resources that relate concepts, preferred terms, variants, definitions, domains, languages, varieties, contexts of use, constraints and evidence.

These structures can intervene during retrieval, classification, generation and evaluation. When necessary, terminology is connected to taxonomies, entities and conceptual models so that understanding, generation and evaluation share a coherent representation of the domain.

Generation conditioned by audience and context

Adapting language to an audience does not simply consist of changing the tone. A system can condition generation on prior knowledge, communicative objective, register, channel, length, depth, language, context and editorial or domain constraints.

The architecture must distinguish between what may vary in the linguistic realisation and what belongs to the factual content and must be preserved.

One and the same knowledge base can therefore produce different explanations for different audiences without needing to invent a different reality for each of them.

Multimodal language: voice, subtitles and video

Language remains subject to constraints when it leaves the page. Voice, subtitles and video introduce duration, rhythm, pronunciation, segmentation, synchronisation and temporal alignment.

We integrate ASR, TTS, linguistic generation, subtitling, alignment and audiovisual composition inside pipelines in which the content can keep a common representation while the medium of realisation changes.

This makes it possible to coordinate text, voice and video without treating each format as an independent process.

Semantic and linguistic evaluation

Evaluating language does not consist solely of measuring textual similarity. Two sentences can use different words and preserve the same meaning; they can also be very similar and contain a critical contradiction.

Depending on the problem, we combine evaluation of semantic fidelity, coverage, contradictions, entailment, terminology, register, naturalness, consistency and preservation of invariants.

We use deterministic validators when a condition can be checked exactly, models when the evaluation requires interpretation, and human review when there is ambiguity, risk or a need for linguistic authority.

A generative evaluator does not replace the deterministic checks the system can carry out directly.

Human-in-the-loop and linguistic governance

Human intervention should not be limited to correcting at the end what a model produces. It should sit where it contributes information, judgement or authority that the system cannot obtain on its own.

That includes linguistic research, the definition of criteria, the resolution of ambiguities, the review of difficult cases, the creation of reference sets, adversarial evaluation, auditing and approval when the context requires it.

The aim is to use specialist review to improve the system and its criteria, not to turn every output into a manual process.

Technologies and methods

Representation and retrieval

Embeddings · rerankers · hybrid retrieval · structured extraction · entity linking · terminologies · taxonomies · conceptual models

Generation

LLMs · hybrid NLG · structured generation · schema-constrained outputs · rules · templates · retrieval-augmented generation · narrative planning

Linguistic engineering and localisation

NLP · linguistic classification · contrastive research · terminology memories · linguistic variation · multi-variety adaptation · synthetic corpora

Evaluation

Semantic similarity · entailment · invariants · criteria-based evaluation · deterministic validators · adversarial evaluation · model-based evaluators · human review

Multimodal

ASR · TTS · temporal alignment · subtitles · synchronisation · audiovisual composition

Language can be generative and still be engineering.

A serious linguistic system is not measured solely by the quality of what it writes. It is also measured by what it represents, by the constraints it preserves, by the transformations it can explain and by our ability to detect when an output has stopped meaning what it was supposed to mean.

We design that complete layer: from knowledge to language, and from language back to a representation that can be checked.

Let's think about how your knowledge should be understood and expressed.

Share the context with us and we will think together about how to combine information, technology, software, people and evaluation around the result that matters.

Search · Type to search. Esc to close.

Search results →

Ask the bot

This is Ask 3.14, an automatic assistant. It answers only with the public content of this website — solutions, capabilities, systems, questions, readings and laboratory work —, shows the content it has used and distinguishes what the website does not yet allow it to establish. It is not a person from the team and it does not know your case; to speak with someone, write to the team.

Automatic assistant limited to the public content of 3.14; it may abstain. To speak with a person, write to the team.

Write to the team →