← Back to Home

Local LLM Tools

Explore the highest-rated open-source tools for running Large Language Models locally. Discover inference engines, GUI clients, and API wrappers for GGUF and safetensors models. Sorted by GitHub authority and active contributions.

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool

Your local-first AI operator. Beats Hermes on GAIA L1 (69.8% vs 58.5%).

Open Dungeon — the first easy-to-use, fully local AI roleplay app. Story and inline scene images

Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device.

Persistent memory + local AI for coding agents. 1.7B–32B open-weight LLM fleet, cross-session Mind

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with

Free macOS-native Cluely alternative. AI chat invisible to screen sharing. Ollama, OpenAI, Claude.

Ask questions about your Obsidian notes using local AI. Privacy-first RAG with Ollama, LM Studio, or

MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works

AI-powered CLI tool: Transform trading research papers into QuantConnect algorithms

Curated real-world use cases for Hermes Agent — the self-improving AI agent from Nous Research.

Extend the Ollama API with dynamic AI tool integration from multiple MCP (Model Context Protocol)

InferrLM - On-device AI for iOS & Android

Fulloch - The Fully Local Home Voice Assistant

Portable multi-agent AI developer setup for Claude Code + Ollama. Role-based local LLM orchestration

A private Claude-Code-style coding agent for Apple Silicon — run chat, code, and local model

Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and

Hyprdots, Hyprland Omarchy, Quickshell, Agentic Local-Ai, Custom Integrated, Minimal

A simple experiment on letting two local LLM have a conversation about anything!

Enterprise-ready self-hosted AI assistant runtime with sandboxed execution, secure credentials,

Modern desktop application (Rust + Tauri v2 + Svelte 5 + Candle (HF)) for communicating with AI

⚡ A local, privacy-focused AI desktop assistant for Windows. Control your PC remotely via Telegram

More than just Karpathy’s LLM Wiki, 100% local with Ollama. Drop Markdown notes → AI extracts

Memory that learns what works.