A generic language model knows the world, but it doesn't know your organisation. With RAG (Retrieval-Augmented Generation) architectures we connect models to your organisation's real data — documents, knowledge bases, management systems — so answers are relevant, up to date and verifiable, with references to the sources they come from. With agents we take the next step: the AI doesn't just answer, it carries out tasks — querying systems, preparing documents, starting workflows — within rules and permissions defined by you.
It's not theory: we've built it
Our on-premise RAG platform ACME ECMS IA & RAG Integration (Private, Simple & Fast RAG) is a fully containerised stack, running on dedicated GPUs in an isolated private network: in our Private Cloud & VDC or directly in the customer's data centre. All the components are open source and the models are interchangeable (Mistral, Gemma, Qwen…): no calls to external services, no data leaving the perimeter, no vendor lock-in.
How it works: the integrated technologies
The logical flow has two paths — ingestion, which turns your documents into a private knowledge base, and query, which takes you from the question to the answer with sources — plus a quality measurement loop.
See it in action
A simulation of the platform's real life cycle: the stack starting up, the components integrating, and a request travelling through the RAG — from the user's shell to the latent space where knowledge is indexed, through to the answer with its sources.
Demo: from docker compose up to the answer with sources
What makes our platform different
- Not just vectors: a knowledge graph. Alongside the vector index, LightRAG builds a graph of entities and relationships (Apache AGE on PostgreSQL): answers connect the facts, rather than merely resembling the question
- Dedicated re-ranking (vLLM): retrieved passages are re-ordered by relevance and filtered before they reach the model — less noise, more precise answers
- A planner that decides: the agent assesses every question and chooses between a direct answer and RAG, explaining why — no pointless searches, no made-up answers
- Quality measured, not claimed: RAGAS generates sets of test questions from your own documents and evaluates the answers with metrics, in Italian
- Images too: the stack includes on-premise image generation (ComfyUI), with the same confidentiality as text
Governed as a service, not as an experiment
- Data never leaves the perimeter: models, indexes and documents live in the dedicated private network — GDPR compliance by design
- Open-source components and interchangeable models: no vendor lock-in, predictable costs (the hardware is yours or on a fixed fee, not billed per token)
- Operations included: container monitoring, administration, persistent volumes and backups with verifiable restore — with optional oversight by our NOC
- Tailored to you: the platform is a proven starting point; models, data sources, permissions and integrations adapt to your context
Typical use cases
- Intelligent document search: query archives, resolutions, manuals and procedures in natural language — with the source cited
- Internal assistants built on your corporate knowledge base: help desk, operational support, onboarding
- Document analysis and preparation: summarising, classification, data extraction from case files and contracts
- Agents in your processes: repetitive tasks carried out by the AI within your workflows, with human approval where needed
Who it's for: public administration, healthcare and businesses that want the benefits of generative AI without giving up confidentiality, compliance and control — or that have already tried "generic" AI and seen its limits on their own data.