Local GenAI Lab

ci license java spring boot react ollama bedrock huggingface mcp rag

Local-first GenAI lab for building and testing tool-assisted AI workflows with local Ollama models or remote provider APIs such as Amazon Bedrock and Hugging Face.

This project goes beyond a chatbot interface by combining a React frontend, a Spring Boot orchestration backend, local and remote model providers, persistent session memory, MCP-backed AWS tooling, and a RAG workspace over the project documentation.

The goal of this project is to explore and demonstrate selected generative AI engineering concepts through a cohesive Java application, rather than to build a feature-complete AI platform.

Local GenAI Lab Agent workspace

Fastest Path

For the shortest path to a running local setup:

cp .env.example .env
./scripts/start.sh

Confirm the local environment file is ignored before adding provider tokens:

git check-ignore -v .env

Then open:

  • frontend: http://localhost:5173
  • backend health: http://localhost:8080/actuator/health

The default provider is Ollama. For a first local run, install Ollama and pull the default model:

ollama pull llama3:8b

Use ./scripts/status.sh or make status to inspect the local runtime. PID files and logs are written under .run/.

Why This Matters

Most LLM demos stop at chat. This project explores how to connect models to real systems while keeping the workflow inspectable on a developer machine.

It demonstrates:

  • backend-side orchestration instead of direct frontend-to-model calls
  • Ollama as the local default, with optional Bedrock and Hugging Face provider APIs
  • persistent local sessions with search, filters, import, and export
  • MCP-backed local tool execution for AWS audits, reports, and artifacts
  • a separate RAG workspace for questions over the local documentation corpus
  • streaming responses, structured report rendering, and API observability

Architecture at a Glance

The frontend has two distinct workspaces: Agent and RAG. They share provider selection and local session storage, but they do not share request orchestration. Agent requests may invoke MCP-backed tools. RAG requests query the local documentation corpus and return cited source chunks.

flowchart LR
  Agent[Agent workspace] --> Backend[Spring Boot backend]
  Rag[RAG workspace] --> Backend

  Backend --> Providers[Ollama / Bedrock / Hugging Face]
  Backend --> Sessions[Local JSON sessions]
  Backend --> Retrieval[Local docs retrieval]
  Retrieval --> Qdrant[Optional Qdrant vector store]

  Backend --> MCP[MCP server]
  MCP --> Scripts[Shell scripts]
  Scripts --> AWS[AWS CLI]
  Scripts --> Reports[Report artifacts]
  Backend --> Reports

For the full architecture, request flows, storage model, and design tradeoffs, see docs/architecture.md and docs/architecture-overview.md.

Main Capabilities

  • Agent workspace for normal chat, streaming chat, tool-assisted prompts, provider selection, sessions, exports, and artifact inspection.
  • Provider abstraction across Ollama, Amazon Bedrock, and Hugging Face, with provider metadata returned in API responses and saved in session history.
  • MCP-backed tool routing for local AWS audit and reporting workflows.
  • Structured report cards and read-only artifact previews under the configured reports directory.
  • Local JSON-backed session storage for chat and RAG conversations.
  • RAG workspace for asking questions against the repository documentation with cited source chunks and saved RAG sessions.

Example Agent Prompts

The Agent workspace can route natural-language AWS questions to local MCP-backed tools when AWS credentials and tool prerequisites are available.

Examples:

  • Analyze my AWS account and summarize the services I am using, highlighting anything unusual or potentially worth reviewing.
  • Analyze my AWS account and generate a summary report of the resources and services currently in use.
  • List my S3 buckets.
  • Run an S3 report for <bucket-name> for the last month.

For the scripts behind these prompts and more examples, see agents/README.md.

Example RAG Prompts

The RAG workspace answers questions against the local docs/ corpus with cited source chunks. It does not run Agent tools or MCP-backed scripts.

Examples:

  • How do I run the project locally?
  • What is the difference between the Agent workspace and the RAG workspace?
  • How does RAG retrieval work in this project?
  • When should I use lexical retrieval versus vector retrieval?
  • How do I troubleshoot Qdrant-backed RAG?
  • What release checks should I run before publishing a release?

Try the same prompt with Lexical, Vector - In Memory, and Vector - Qdrant to compare retrieved sources and answer quality.

RAG Workspace

The RAG workspace is isolated from the normal Agent flow. It does not invoke MCP tools or agent routing. It loads the local docs/ corpus, retrieves relevant chunks, and asks the selected provider to answer with citations.

Local GenAI Lab RAG workspace

Current RAG support includes:

  • lexical in-memory retrieval by default
  • optional in-memory vector retrieval
  • optional Qdrant-backed vector retrieval
  • retrieval target controls and technical timing details in the UI
  • local RAG session persistence with answers and citations

RAG details live in:

Run Locally

Prerequisites:

  • Java 21
  • Maven 3.9 or newer
  • Node.js 20.19 or newer
  • Ollama for the default local provider path
  • Docker and Docker Compose for Docker validation or optional Qdrant; Trivy for image scanning
  • AWS CLI, jq, and AWS credentials only for AWS tool flows

Common commands:

make help
make start
make status
make local-verify
make test
make verify
make release-check

For remote Linux or EC2 development, make local-verify is the most explicit local validation entry point. It checks the Java/Maven/Node/npm toolchain, runs the supported verification flow, and writes long command output to /tmp/local-genai-lab-*.txt so failures are easier to inspect over SSH.

Docker-inclusive release validation is available when Docker and Trivy are installed:

make release-check-docker

On Amazon Linux / EC2 hosts, Trivy is commonly missing by default. One working install pattern is:

sudo rpm -ivh https://github.com/aquasecurity/trivy/releases/latest/download/trivy_0.66.0_Linux-64bit.rpm
trivy --version

For Docker-based AWS Agent tools, copy .env.docker-aws-tools.example to .env.docker-aws-tools. The file is ignored by Git and lets the normal Docker scripts mount your local AWS configuration read-only into the backend container. After Docker starts, verify the mounted identity before Agent testing:

./scripts/docker-aws-preflight.sh

Docker Agent Testing

Use the complete preparation workflow rather than restarting containers alone:

./scripts/docker-go.sh

It builds the current source, restarts Docker, smoke-checks the running stack, and verifies that the backend container can authenticate to AWS. Use ./scripts/docker-go.sh --skip-build only when intentionally validating the existing Docker image and configuration without a fresh local build.

docker-go.sh works on macOS and Linux wherever Bash and Docker Compose are available. On Windows 11, run it from WSL or Git Bash rather than native PowerShell.

Where to test:

  1. If Docker runs locally on your Mac or Windows computer, run docker-go.sh there and open http://localhost:3000. Do not create an SSH tunnel.
  2. If Docker runs on EC2 or another remote host, run docker-go.sh on that remote host. Then create an SSH tunnel from your Mac or workstation using a separate local port:
ssh -N -L 3001:localhost:3000 my-ec2-1

Leave the tunnel open and test the remote deployment at http://localhost:3001 from your Mac or workstation.

Replace my-ec2-1 with an SSH alias from the workstation's ~/.ssh/config, or with a full SSH destination such as [email protected]. After frontend changes, use an Incognito window or DevTools Empty Cache and Hard Reload before testing.

Provider setup details are in docs/providers.md. Testing and release validation details are in docs/testing.md and docs/release-checklist.md.

Project Structure

local-genai-lab/
|-- backend/      Spring Boot API, provider orchestration, sessions, RAG APIs
|-- frontend/     React UI for Agent and RAG workspaces
|-- agents/       MCP-facing shell tools and generated report artifacts
|-- mcp/          local MCP server
|-- scripts/      lifecycle, build, Docker, and release scripts developers run directly
|-- ops/          local smoke checks and shell test support
|-- docs/         architecture, testing, troubleshooting, RAG, and ADR docs
|-- data/         local runtime data, including JSON session storage
`-- Makefile      command entry points for local development and validation

Component details:

Documentation Map

Start here:

RAG and retrieval:

Project governance and release references:

Current Scope

  • single-user, local-first GenAI lab
  • built for hands-on learning, AWS Generative AI Developer Professional exam preparation, and technical experimentation
  • optimized for correctness, inspectability, and local workflow clarity rather than multi-user scale
  • intended to run on a developer machine with local Ollama, optional Bedrock access, and optional MCP-backed AWS tooling

Known Limitations

  • not designed as a multi-tenant or internet-facing production service
  • MCP tool execution uses short-lived local subprocesses
  • backend health/readiness is backend-only; whole-stack checks belong to ops/check-app.sh
  • artifact access is intentionally read-only and bounded to the configured reports directory
  • Bedrock and AWS tool flows depend on local AWS credentials and runtime setup
  • larger local models can be slower and may require higher backend read timeouts

Contact

License

This project is licensed under the MIT License.