Skip to main content

Artificial intelligence

Internal document assistant with guardrails and evaluation

Augmented retrieval over a corpus of standards and technical notes

Cabinet de conseil technique (exemple) - 2025-02 - 2025-06

Demonstration sample: this case study illustrates the Codexway method and does not refer to a real client.

Context

A technical consultancy wanted to reduce the time its consultants spent searching for references across internal notes, standards and engagement reports. The corpus held around eleven thousand heterogeneous documents, with multiple versions and different access rights per team.

Problem

The existing search relied on keyword matching, which performed poorly on business wording. Consultants often reconstructed answers from memory, risking a citation of an outdated standard version or an unsourced recommendation.

Objectives

  • Provide answers grounded in identified, cited passages
  • Respect existing document access rights
  • Measure quality objectively before any generalisation

Role

Owner of the design and implementation, working with the internal quality lead to build the business evaluation set.

Solution

The project began with the corpus: source inventory, removal of outdated versions, verification of usage rights and chunking into passages sized to the questions asked. This step accounted for nearly half of the total effort and revealed that eleven percent of documents were duplicates or superseded versions. The retrieval chain combines lexical and vector search, with prior filtering on the user's access rights. Every answer cites the passages used, with a link to the source document and its version date. When retrieval returns nothing relevant, the assistant says so explicitly instead of producing an approximate answer. Evaluation relies on a set of one hundred and twenty representative questions, with expected answers and legitimate sources. This set is replayed on every change to the chain, which surfaces quality regressions before release.

Architecture

A Spring Boot service built on Spring AI for orchestration, vector search and model calls. Documents and their passages are stored in PostgreSQL with the pgvector extension; indexes are rebuilt by a batch Python process. Access rights are resolved from the existing directory and applied as a filter before retrieval. Redis holds recent conversation context and limits the cost of redundant calls. The whole system runs on Kubernetes, with logging of questions, answers and consulted sources.

Results (samples)

  • Cleaned and indexed corpus with access rights tracking
  • Answers systematically grounded in cited passages
  • Business evaluation set replayed on every change
  • Explicit no-answer behaviour when the corpus lacks the information

Indicators (samples)

  • Documents processed and indexed

    around 11,000

    Simulated corpus for the demonstration — demonstration sample

  • Questions in the evaluation set

    120

    Illustrative evaluation set — demonstration sample

  • Documents identified as duplicates or outdated

    11 percent

    Simulated cleanup result — demonstration sample

Technologies

  • Python
  • Spring AI
  • PostgreSQL
  • pgvector
  • Redis
  • Kubernetes
  • API de modèles de langage

Testimonial (sample)

« What changed how people work is the systematic citation of sources. Our consultants check the document version before reusing a recommendation. »

- Technical consultancy (demonstration sample)

Related services

  • Application security and API protection

    I secure your applications where it actually matters: identities, authorisations, secrets and input validation. The measures are verifiable and integrated into the delivery pipeline.

  • Applied generative artificial intelligence

    I design useful, measurable generative AI use cases: document search, summarisation, writing assistance and controlled answers. Every feature is evaluated on real cases before being exposed to users.

  • Process automation with AI agents

    I automate repetitive processing by combining deterministic rules with language models and explicit guardrails. Automation handles the simple cases and hands ambiguous ones to a human.

  • Observability and application performance

    I make your applications diagnosable: correlated traces, useful metrics and alerts that signal a real problem. You move from interpreting symptoms to identifying the cause.