Privacy
Keep your AI workflows closer to your own machine.
Run powerful language models on your machine with a polished desktop experience, live performance monitoring, flexible model support, and private local workflows.
View on GitHubWhy local
Cloud AI is convenient, but local AI gives you more direct control over your data, your models, your runtime, and how your hardware is used.
Keep your AI workflows closer to your own machine.
Choose the model, runtime, and configuration that fits your workflow.
See how your hardware is performing while the model runs.
Experiment with compatible local models without being locked into one provider.
Designed for local-first workflows.
Features
An intuitive conversational workspace designed for local models.
Work with the models you actually want to run.
Know what your machine is doing.
Organize, rename, pin, and remove conversations in the app.
Control how your local AI environment behaves.
A polished desktop application built for Windows.
Find and download compatible models directly in OffyAI through its connection to the public Hugging Face API.
Product tour
http://localhost:8080Dark4096EnabledPerformance monitoring
OffyAI exposes useful runtime information while AI workloads are running, so you can see how your hardware responds in real time.
Metrics shown above are illustrative UI examples, not live statistics from this website.
Model support
OffyAI includes a protected default model named offyai.gguf and lets you find and download compatible models through its library connected to the public Hugging Face API.
How it works
Select a compatible model from the library or place one in the local models directory.
OffyAI connects to the configured llama.cpp-compatible local backend.
Choose the model and interact through the desktop interface.
Track performance, manage conversations, and configure the environment.
Privacy
OffyAI is designed to keep local inference and app data on your machine, subject to the model, runtime, and external services you choose to configure.
Local model execution
Local conversation sessions
Local configuration
No mandatory cloud inference
User-controlled model files
Optional external integrations
API integrations, updates, downloads, or other configured external services may involve network communication.
Architecture
OffyAI application stack
System requirements
Windows 10 / 11
Enough space for the application and any selected models
Depends on model size and quantization
Optional, depending on runtime configuration
Modern 64-bit processor recommended
A compatible local GGUF, BIN, or GGML model
Large models require substantially more memory and compute resources.
Download
Checking for the latest release…
What's new
Release notes will appear here once fetched from GitHub.
Installation
.exe file.FAQ
OffyAI is a Windows desktop application that provides a polished, ChatGPT-like interface for running compatible language models on your own machine, alongside model management and performance monitoring tools.
Yes. OffyAI is designed for local-first workflows, connecting to a llama.cpp-compatible runtime running on your machine.
Core chat with a local model does not require a mandatory cloud dependency. Downloads, updates, and any external services you choose to configure may still require network access.
No. OffyAI provides a ChatGPT-like interface but runs compatible local language models through a llama.cpp-compatible runtime rather than depending on ChatGPT itself.
GGUF, BIN, and GGML, subject to llama.cpp compatibility.
Yes. OffyAI includes a model library connected to the public Hugging Face API, so you can find and download compatible models for local use.
GGUF is a model file format commonly used with llama.cpp-based local inference runtimes.
Yes. OffyAI ships with a built-in default local model file named offyai.gguf, which is the initial model used when the app starts.
Conversation sessions are managed in the desktop app, with storage behavior determined by the app and runtime environment.
Application settings, including theme, server URL, and model configuration, are stored locally.
Yes. Place a compatible GGUF, BIN, or GGML model in the local models directory and select it in the application.
GPU usage depends on your runtime configuration and hardware; GPU monitoring is shown when available.
Requirements depend on the model you choose. See the System Requirements section for general guidance.
Check the project's GitHub repository or release information for current licensing terms.
Report bugs through the GitHub Issues page for this repository.
Feature requests can also be submitted through GitHub Issues.
Use the Download section on this page, or visit the GitHub Releases page directly.
Help
Verify the model path and confirm the file is a supported format (GGUF, BIN, or GGML).
OffyAI validates available local model files; an invalid or unreadable file may not be listed.
Speed depends on model size, quantization, CPU, GPU, RAM, and runtime configuration.
Larger models require more memory; consider a smaller or more quantized model.
Verify the server URL, confirm the runtime is running, and check the port configuration.
Try reinstalling, verify you downloaded from the official source, check Windows compatibility, and review release notes or open issues.
Open source
Explore local models, monitor performance, and build an AI workflow around your own hardware.