---
title: "Building a company GPT with RAG and workflow automation"
description: "How small and midsize companies build their own Company GPT. With RAG system, n8n workflows, and GDPR-compliant infrastructure for your own AI assistant."
canonical: "https://simo-online.com/en/blog/company-gpt-bauen-kmu-leitfaden-2026-mit-rag-n8n"
---

# Building a company GPT with RAG and workflow automation

How small and midsize companies build their own Company GPT. With RAG system, n8n workflows, and GDPR-compliant infrastructure for your own AI assistant.

- Author: SIMO GmbH
- Published: 2025-06-18
- Updated: 2026-10-07
- Topic: [AI Tools & Technology](https://simo-online.com/en/blog/topic/ai-tools-technology)

A Company GPT of your own is no longer a dream of the future but a concretely achievable solution for small and midsize companies. What this means is an AI system based on your own company knowledge that answers employee questions and can be integrated into existing business processes—without sensitive data having to flow to external providers.

This guide shows step by step how midsize companies build their own AI assistant, which technologies are needed, and what to consider regarding data protection and governance.

## What a company GPT means for midsize companies

A Company GPT is essentially a language model linked to internal company knowledge. Instead of drawing on general internet knowledge, it responds based on your documents, process descriptions, price lists, and customer information.

The advantages are obvious:

- Instant Answers: Employees ask questions in natural language and receive precise answers with source citations.
- Knowledge Preservation: The expertise of experienced employees is systematically captured and made accessible.
- Faster Onboarding: New team members find their way more quickly because they have access to an intelligent knowledge database.
- Process Efficiency: Recurring questions no longer consume human work time.

## The technical architecture: RAG as the foundation

The heart of a Company GPT is a RAG system (Retrieval Augmented Generation). RAG combines two strengths: the ability of a language model to understand and generate natural language, and the precision of a vector database that finds relevant documents.

## How RAG works

- Ingestion: Company documents are read in, divided into sections (chunks), and stored as mathematical representations (embeddings) in a vector database.
- Retrieval: When a user asks a question, it is also converted into an embedding. The vector database finds the most semantically similar text sections.
- Generation: The found text sections are passed to the language model together with the question, which formulates a precise answer from them.

## The technology stack

The following stack has proven effective for midsize companies:

- Component: Workflow Engine | Recommended Solution: n8n | Task: Orchestration of all processes
- Component: Vector Database | Recommended Solution: Supabase with pgvector | Task: Storage of embeddings
- Component: Embedding Model | Recommended Solution: OpenAI text-embedding-3-small | Task: Converting text to vectors
- Component: Language Model | Recommended Solution: current language models from OpenAI, Anthropic, or Meta (Llama) | Task: Answer generation
- Component: Re-Ranker | Recommended Solution: Cohere Rerank | Task: Quality optimization of search

This stack is cost-efficient, scalable, and does not require its own IT department.

## Step-by-step setup

## Step 1: identify data sources

Before you begin implementation, take inventory of your company knowledge:

- Structured Documents: SOPs, manuals, price lists, contracts
- Unstructured Documents: Emails, meeting notes, wiki entries
- Systems: CRM data, ERP information, ticketing systems

Start with the 20 to 50 most important documents. A local folder or a Google Drive is perfectly sufficient as a starting point.

## Step 2: document preparation

The quality of your Company GPT stands and falls with the quality of prepared documents:

- PDF Extraction: PDFs with a text layer are unproblematic. Scanned documents require OCR.
- Cleaning: Remove headers, footers, signatures, and duplicate content.
- Metadata: Each document should be tagged with mandatory fields: department, document type, validity period, source.
- Chunking: Split texts into meaningful sections of 512 to 800 tokens, with 10 to 20 percent overlap.

## Step 3: build the ingestion workflow in n8n

The first n8n workflow connects your data sources to the vector database:

- Trigger: Scheduled (daily) or event-based (on file update)
- Read data source: Google Drive, SharePoint, or local folder
- Process documents: Extraction, cleaning, chunking
- Generate embeddings: Call OpenAI API
- Store in Supabase: Chunks with vectors and metadata

## Step 4: create the Q&A workflow

The second workflow processes user queries:

- User asks a question (via chat interface, Slack, or email)
- Question is converted into an embedding
- Vector search in Supabase with metadata filters
- Re-ranking of results with Cohere
- Language model generates an answer based on the top matches
- Answer with source citations is returned

## Step 5: access rights and governance

- Row Level Security (RLS) in Supabase ensures that users can only access authorized documents.
- Restrict the system prompt: The model may only respond based on the delivered documents. When information is missing, it must communicate this transparently.
- Capture audit logging: Who asked which question when? Plan for this from the start in compliance-sensitive industries.

## Practical examples from midsize companies

## Skilled trades business with 18 employees

A sheet metal shop fed its Company GPT with maintenance instructions, supplier terms, and measurement forms. New employees now get their answers in seconds instead of minutes. The system has been running stably for months without notable maintenance.

## Timber construction company with 22 employees

A timber construction firm uses its Company GPT for automatic classification of incoming inquiries, proposal preparation, and as an internal knowledge base for construction plans and supplier lists.

## B2B service provider with 25 employees

A service company uses its Company GPT for lead qualification: The system analyzes lead data from CRM and website, creates automatic scorings, and prioritizes inquiries. The conversion rate rose from 8 to 12 percent.

## Costs and ROI

## Ongoing costs

For a typical midsize-company setup with OpenAI embeddings, Supabase Pro, and n8n Cloud, the monthly costs are in the low double-digit euro range. Costs scale with document volume and the number of queries.

## One-time costs

- Document preparation: The largest initial effort. Depending on document volume, 1 to 5 working days.
- Workflow setup: 1 to 2 days for the prototype, 2 to 4 weeks for a production-ready system.
- Training: AI literacy training for employees (measures under Article 4 have been mandatory since February 2025).

## Typical ROI

- Use Case: Internal knowledge base | Typical Annual Savings: €20,000–€40,000 | Typical Year 1 ROI: 200%–400%
- Use Case: Automatic inquiry classification | Typical Annual Savings: €30,000–€50,000 | Typical Year 1 ROI: 300%–500%
- Use Case: Lead qualification | Typical Annual Savings: €50,000–€180,000 | Typical Year 1 ROI: 400%–650%
- Use Case: AI customer service | Typical Annual Savings: €40,000–€60,000 | Typical Year 1 ROI: 100%–200%

## GDPR and EU AI Act

## GDPR requirements

- Data Processing Agreements (DPA): Conclude with all AI providers (OpenAI offers a DPA).
- Data Minimization: No personal data in the knowledge base unless absolutely necessary.
- Access Control: RLS in Supabase, explicit filters in all queries.

## EU AI Act

- AI literacy: The rule has applied since February 2025. Since July 2026, Article 4 requires measures that support your staff’s AI literacy; document training sessions and guidelines.
- Transparency obligation: In effect since August 2, 2026. Document which data flows through the system.
- Risk Classification: An internal knowledge system is typically not classified as a high-risk system but still requires documentation.

## The 90-day plan for your own company GPT

## Phase 1 (day 1–30): foundation

- Conduct process and data audit
- Identify and prepare 20–50 core documents
- Set up n8n and Supabase
- Define metadata standards
- Hold AI literacy training

## Phase 2 (day 31–60): prototype

- Build ingestion workflow and index first documents
- Create Q&A workflow with simple chat interface
- Create a Golden Set with 10–30 test questions and evaluate
- Configure access rights

## Phase 3 (day 61–90): production

- Connect system to additional data sources
- Integrate re-ranking
- Activate audit logging
- Roll out to all relevant employees
- Start ROI measurement

## Frequently asked questions

Do I need an IT department for a Company GPT?

No. With n8n as a workflow tool and Supabase as a managed database, a functional system can be built without programming skills. The tools are designed so that technically proficient business users can operate them.

How do I keep my Company GPT up to date?

The ingestion workflow can run on a schedule (daily, weekly) or event-driven (on file updates). This way, the knowledge base automatically stays current.

What happens if the system gives a wrong answer?

A correctly configured system prompt forces the model to respond only based on the delivered documents. When information is missing, it responds with “I do not have verified information on that.” Additionally, a Golden Set evaluation process ensures systematic quality assurance.

How do I protect confidential information?

Row Level Security in Supabase, metadata-based filtering, and clear access rights ensure that each user only sees the documents intended for them.

## References

- [OpenAI - Embedding Models and API Updates](https://openai.com/index/new-embedding-models-and-api-updates/)
- [OpenAI - Data Processing Addendum (DPA)](https://openai.com/de-DE/policies/data-processing-addendum/)
- [Supabase - Row Level Security Documentation](https://supabase.com/docs/learn/auth-deep-dive/auth-row-level-security)
- [pgvector - Vector Database Extension for PostgreSQL](https://github.com/pgvector/pgvector)
- [Cohere - Rerank Documentation](https://docs.cohere.com/release-notes/)
- [n8n - Company Knowledge Base Agent (RAG) Workflow](https://n8n.io/workflows/6538-company-knowledge-base-agent-rag/)
- [EU AI Act - Regulatory Framework](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai)
- [ifo Institute - Unternehmen setzen zunehmend auf KI [in German] (2025)](https://www.ifo.de/en/press-release/2025-06-16/companies-germany-increasingly-relying-artificial-intelligence)

Rendered version: https://simo-online.com/en/blog/company-gpt-bauen-kmu-leitfaden-2026-mit-rag-n8n
