Checking latest release…

Local AI, Built for Control and Performance.

Run powerful language models on your machine with a polished desktop experience, live performance monitoring, flexible model support, and private local workflows.

View on GitHub
LOCAL-FIRST PRIVATE GGUF · BIN · GGML DOCUMENT CONTEXT WINDOWS DESKTOP LLAMA.CPP

Why local

AI that stays closer to you.

Cloud AI is convenient, but local AI gives you more direct control over your data, your models, your runtime, and how your hardware is used.

Privacy

Keep your AI workflows closer to your own machine.

Control

Choose the model, runtime, and configuration that fits your workflow.

Performance

See how your hardware is performing while the model runs.

Flexibility

Experiment with compatible local models without being locked into one provider.

Designed for local-first workflows.

Features

Everything you need for local AI.

01

AI Chat

An intuitive conversational workspace designed for local models.

  • Multi-session chats
  • Responsive interface
  • Streaming responses
  • New conversations
02

Model Management

Work with the models you actually want to run.

  • GGUF · BIN · GGML
  • Model discovery
  • Model selection
  • Validation and model discovery
03

Live Performance

Know what your machine is doing.

  • CPU · RAM · GPU
  • Tokens per second
  • Temperature data
  • Performance charts
04

Local sessions

Organize, rename, pin, and remove conversations in the app.

  • Session switching
  • Rename and pin
  • Delete conversations
  • Multiple conversations
05

Configuration

Control how your local AI environment behaves.

  • Local server URL
  • Sampling controls
  • Theme and preferences
  • Runtime tuning
06

Desktop Experience

A polished desktop application built for Windows.

  • Electron-based
  • Bundled llama-server
  • Windows installer
  • Desktop workflow
07

Model Library

Find and download compatible models directly in OffyAI through its connection to the public Hugging Face API.

  • Hugging Face API
  • Model discovery
  • In-app downloads
  • Compatible model files

Product tour

See OffyAI from the inside.

Ready. Model loaded and listening on the local runtime.
Explain what changed in this diff.
The file context is attached to the conversation and the local model is ready to respond.
Message OffyAI…

Performance monitoring

Don't just run a model. Understand it.

OffyAI exposes useful runtime information while AI workloads are running, so you can see how your hardware responds in real time.

CPU 42%
Memory 6.8 GB
GPU 37%
Generation 28.4 tok/s
Response 1.8s

Metrics shown above are illustrative UI examples, not live statistics from this website.

Model support

Built-in model support.

OffyAI includes a protected default model named offyai.gguf and lets you find and download compatible models through its library connected to the public Hugging Face API.

GGUF Local runtime Default model
  • Built-in offyai.gguf model
  • Local inference workflow
  • Model upload and selection
  • Model library connected to the public Hugging Face API
  • Download compatible models in OffyAI
Load Chat Monitor

How it works

From install to inference.

  1. 01

    Choose or download a model

    Select a compatible model from the library or place one in the local models directory.

  2. 02

    Start the local runtime

    OffyAI connects to the configured llama.cpp-compatible local backend.

  3. 03

    Start chatting

    Choose the model and interact through the desktop interface.

  4. 04

    Monitor & manage

    Track performance, manage conversations, and configure the environment.

Privacy

Your machine. Your models. Your control.

OffyAI is designed to keep local inference and app data on your machine, subject to the model, runtime, and external services you choose to configure.

Local model execution

Local conversation sessions

Local configuration

No mandatory cloud inference

User-controlled model files

Optional external integrations

API integrations, updates, downloads, or other configured external services may involve network communication.

Architecture

Under the hood.

OffyAI application stack

Electron
OffyAI Desktop UI
Local Application Logic
llama.cpp-compatible Runtime
Local Model
User Hardware
Electron Next.js React Node.js llama.cpp GGUF

System requirements

What you'll need.

Operating System

Windows 10 / 11

Storage

Enough space for the application and any selected models

Memory

Depends on model size and quantization

GPU

Optional, depending on runtime configuration

CPU

Modern 64-bit processor recommended

Model

A compatible local GGUF, BIN, or GGML model

Large models require substantially more memory and compute resources.

Download

Run OffyAI on your Windows machine.

Checking for the latest release…

Latest version
Release date
Installer name
Installer size

What's new

Latest release notes.

Release information unavailable

Release notes will appear here once fetched from GitHub.

View full release notes on GitHub

Installation

Get up and running.

  1. Download the Windows installer.
  2. Run the .exe file.
  3. Install OffyAI.
  4. Place compatible model files in the models directory.
  5. Configure or select the model.
  6. Start chatting.
  7. Monitor performance.
Security note: if Windows displays a security warning, verify that you downloaded the installer from the official GitHub release before proceeding.

FAQ

Common questions.

Help

Troubleshooting.

Model not found

Verify the model path and confirm the file is a supported format (GGUF, BIN, or GGML).

Model does not appear

OffyAI validates available local model files; an invalid or unreadable file may not be listed.

Slow generation

Speed depends on model size, quantization, CPU, GPU, RAM, and runtime configuration.

High memory usage

Larger models require more memory; consider a smaller or more quantized model.

Local server not connecting

Verify the server URL, confirm the runtime is running, and check the port configuration.

Application does not start

Try reinstalling, verify you downloaded from the official source, check Windows compatibility, and review release notes or open issues.

Open source

Built in the open.

Stars
Forks
Open issues

Take AI back to your machine.

Explore local models, monitor performance, and build an AI workflow around your own hardware.