Bookmarks

Top 10 GitHub Repos — Full Archive

CLI

Kimi Code CLI
An AI coding agent that runs in your terminal — it can read and edit code, run shell commands, search files, fetch web pages, and choose the next step based on the feedback it receives. Works out of the box with Moonshot AI's Kimi models and can be configured to use other providers. A fresh, terminal-native coding agent from a major model lab — a new default option among open-source agent CLIs.
esengine/DeepSeek-Reasonix
A DeepSeek-native AI coding agent for the terminal, shipped as a single MIT-licensed Go binary and engineered around prefix-cache stability so you can leave it running. One engine exposes four surfaces — terminal, desktop app, browser, or your editor over ACP — with plan mode, permissions, and a workspace sandbox. Extensions and spec docs are included. A new single-binary coding agent built for long-running stability and four-way ACP access.
OpenAI Codex
A lightweight coding agent from OpenAI that runs locally in your terminal, with IDE integrations for VS Code, Cursor, and Windsurf, plus a desktop app path and the cloud-based Codex Web. It remains the reference point for terminal-native coding agents and a backbone many agent tools build on or integrate with. Still the most widely adopted terminal coding agent — essential context for anyone building agent tooling.
Entire CLI
Entire CLI hooks into git and records what an AI coding agent actually did during a session — the prompts, the files touched, the tool calls — then indexes that transcript alongside the commit it produced. You can resume a session from any earlier checkpoint, or hand a teammate the full "why did the code change" story instead of just a diff. It shipped v0.10.2 on August 19, six days before this list went out, and has already logged nearly 8,000 commits of its own. Turns agent sessions into a searchable part of git history instead of a conversation that evaporates.
vercel-labs/skills
The CLI for the open agent skills ecosystem. Supports OpenCode, Claude Code, Codex, Cursor, and 68 more agents. Install skills with `npx skills add` and use them interactively or pipe into any supported coding agent. A single command to discover and invoke agent skills across the ecosystem. Standardizes the "agent skills" concept across 70+ agents — Vercel is building the package manager for agent capabilities.
google/agents-cli
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud's Gemini Enterprise Agent Platform. Includes pre-built skills for agent creation, evaluation, and deployment workflows. Google's official CLI for agent lifecycle management on GCP — signals enterprise-grade agent tooling from a major cloud provider.
agent-device
A CLI that lets AI agents drive real iOS and Android devices/emulators — ADB, simulators, XCUITest, Appium-style automation exposed for agentic control and E2E testing. Created 2026-01-30, it fills an obvious gap: most agent tooling targets the web or the shell, leaving mobile largely untouched. A clean primitive for agent-driven mobile QA and automation. Brings mobile devices into reach for agents — a new modality (mobile control/testing) most harnesses can't touch today.
terax-ai
A 7MB, terminal-first "AI-native dev workspace" (Tauri + Rust) created on 2026-04-21. It targets developers who want an editor/agent surface that lives in the terminal without the bloat of a full IDE, with cross-platform builds for Linux, macOS, and Windows. The combination of tiny footprint and fast star growth makes it a notable new entrant in the agentic-CLI space. Lightweight, terminal-native agent workspace gaining traction fast.
rmux
A universal Rust multiplexer with a typed SDK to drive any CLI or TUI app from code — native on Linux, macOS, and Windows. Created 2026-05-15, it's effectively a programmable harness for automating terminal programs, which is increasingly relevant for agents that need to operate interactive command-line tools reliably. The youngest repo here, and a quietly important building block for terminal automation. The newest pick — a clean primitive for letting agents drive interactive CLI/TUI apps.
farion1231/cc-switch
A cross-platform desktop all-in-one assistant for Claude Code, Codex, OpenCode, Gemini CLI, and Hermes Agent. Manages multiple coding agents from a single interface. Consolidates the growing ecosystem of coding agents into one manager — a sign of the maturing toolchain.
Andyyyy64/whichllm
Find the local LLM that actually runs and performs best on your hardware. Auto-detects GPU/CPU/RAM and ranks top models from HuggingFace that fit your system. One command, run it instantly, ranked by real, recency-aware benchmarks. Solves the "which model fits my hardware?" problem with a single CLI command — essential for anyone running local LLMs.
graykode/abtop
Like `htop`, but for AI coding agents. Monitors Claude Code, Codex CLI, and OpenCode sessions in real-time — token usage, context window percentage, rate limits, child processes, open ports, and more. Essential ops tool for anyone running multiple coding agents; fills the same niche `htop` does for system processes.
kunchenguid/no-mistakes
A local git proxy that intercepts pushes and runs an AI-driven validation pipeline before forwarding to the real remote. Spins up a disposable worktree, validates the branch, and opens a clean PR automatically — but only after every check passes. Designed to "kill all the slop" and raise clean PRs. Addresses the real pain point of AI-generated code quality gatekeeping in a clever, non-blocking way.
modem-dev/hunk
A review-first terminal diff viewer built specifically for agent-authored changesets. Features multi-file review with sidebar navigation, inline AI annotations, split/stack layouts, and watch mode. Built on OpenTUI and Pierre diffs. Agent-generated code needs agent-aware review tools — this fills a UX gap that traditional diff viewers don't address.
whale
A Go-based DeepSeek-native terminal coding agent that ships MCP support, agent skills, shell tools, and prefix-cache optimisation out of the box — effectively a DeepSeek-first answer to Claude Code's CLI. Created 2026-05-06; small but well-scoped and uses the modern coding-agent feature set rather than reinventing it. First credible *terminal-native* coding agent built around DeepSeek as the primary model rather than as an afterthought provider.
token-tracker
A token-usage and cost dashboard for local AI agents — custom Claude Code StatusLine, CLI dashboard with cost analysis, rate-limit monitoring, and session tracking. Python with the `rich` terminal UI, 23 forks, MIT-style permissive. Created 2026-05-08. Sits next to `caveman` as the second half of the "what are my agents actually costing me?" answer — `caveman` cuts tokens, `token-tracker` shows the bill. The first credible cost-observability layer specifically for Claude Code + Codex stations, before paid SaaS dashboards catch up.
cli (Google Workspace)
Google's official one-command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, and Admin — dynamically built from the Discovery Service and shipping with AI-agent skills baked in. Created 2026-03-02, it's a first-party signal that vendors are now designing CLIs explicitly for agents to drive, not just humans. First-party Google tooling built for agents — a template for vendor CLIs in the agent era.

MCP

MCP TypeScript SDK
The official TypeScript SDK for Model Context Protocol servers and clients, now on the v2 line (`@modelcontextprotocol/server`, `@modelcontextprotocol/client`) implementing the 2026-07-28 MCP spec. The v2 rework addresses long-standing architectural issues, and the docs start with a ten-minute server tutorial. The reference SDK for the current MCP spec — relevant to anyone shipping MCP servers or clients in TypeScript.
jCodeMunch MCP
An MCP server for precise, symbol-level GitHub source retrieval via tree-sitter AST parsing, compatible with Claude Code, Cursor, and any MCP client. Claims 95%+ token-cost reduction on code exploration, with 313B+ tokens already saved. Attacks the biggest practical cost in agentic coding — burning tokens on code navigation — at the retrieval layer.
modelcontextprotocol/rust-sdk
The official Rust implementation of the Model Context Protocol, built on tokio with an async runtime. The repo houses the core `rmcp` crate plus `rmcp-macros` for procedural macro support, and the 3.x migration guide shows the SDK actively tracking spec evolution. For teams building MCP servers or clients in Rust. The official SDK matters to any Rust-based MCP work, and the 3.x line is actively moving.
different-ai/openwork
A free, open-source desktop app for sharing AI workflows across agents on macOS, Windows, and Linux — positioned as an open alternative to Claude Cowork, powered by opencode. Add a single OpenWork MCP to Codex, Claude Code, Cursor, or another compatible agent and reuse the same skills, MCPs, and connected services across tools and team. One MCP to rule the workflow layer. Shows the MCP pattern applied as a cross-tool workflow-sharing layer, not just a chat-client integration.
gortex
A high-performance, graph-based code-intelligence engine for AI agents and IDEs supporting 257 languages and multi-repository setups, exposed via CLI, MCP Server, and API. It aims to expose only the information an agent needs, claiming token usage reductions of up to 50x — all fully local. MCP-native code intelligence that directly attacks the token-cost problem for AI coding agents.
ChromeDevTools/chrome-devtools-mcp
An official MCP server from Chrome DevTools that lets coding agents (Antigravity, Claude, Cursor, Copilot) control and inspect a live Chrome browser. Gives AI assistants access to the full power of Chrome DevTools for reliable automation, in-depth debugging, and browser inspection. First-party MCP server from Chrome DevTools — gives every coding agent browser automation and debugging capabilities without third-party wrappers.
DeusData/codebase-memory-mcp
High-performance code intelligence MCP server that indexes codebases into a persistent knowledge graph. Full-indexes an average repo in milliseconds (Linux kernel at 28M LOC in seconds). Supports 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies. Dramatically reduces token usage for agent code understanding — the speed and efficiency numbers make this a must-try for any agent-heavy workflow.
wonderwhy-er/DesktopCommanderMCP
MCP server for Claude that provides terminal control, file system search, and diff file editing. Goes beyond typical AI editors by running processes, managing files, and automating tasks — all while using host client subscriptions instead of API token costs. One of the most practical MCP servers for developers who want Claude to operate their full development environment.
semble
Fast, accurate semantic code search for agents, shipped as an MCP server — and pitched as using ~98% fewer tokens than the usual grep-then-read loop. Created 2026-04-06, it's climbing fast because token-efficient retrieval is one of the highest-leverage optimizations for any coding agent, and it drops straight into Claude Code / Cursor via MCP. Token-frugal code search over MCP — a direct fix for the most expensive habit coding agents have.
awesome-mcp-servers
The most-referenced catalog of Model Context Protocol servers — a continuously updated index of connectors spanning databases, browsers, cloud services, and developer tools. As MCP becomes the lingua franca for tool access, this list is the practical starting point for anyone wiring servers into Claude, Cursor, or a custom client. The ~1,400 open issues are largely submission/PR traffic, a signal of how fast the ecosystem is still growing. The single best map of the MCP server landscape, kept current by the community.
design-extract
Extracts a website's complete design system in one command — DTCG tokens, semantic/primitive/composite layers — and exposes it as an MCP server for Claude Code, Cursor, and Windsurf. It emits to multiple platforms (SwiftUI, Compose, Flutter, Tailwind v4, Figma variables, shadcn/ui) and includes a CSS health audit and WCAG remediation. Created 2026-04-15, it's a sharp example of a focused, practical MCP tool rather than a general framework. A genuinely useful, narrowly-scoped MCP server for design-to-code workflows.
chopratejas/headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM — achieving 60-95% fewer tokens with the same answers. Available as a library, proxy, and MCP server. Directly addresses the token cost problem for agent tool calls and RAG pipelines via an MCP-native approach.
aws/agent-toolkit-for-aws
Official AWS-supported MCP servers, skills, and plugins that help AI coding agents build, deploy, and manage applications on AWS. Works with Claude Code, Codex, Cursor, and other major coding agents. Provides the tooling, knowledge, and guardrails for agentic AWS development. First-party AWS MCP support signals enterprise readiness for agent-driven cloud workflows.
serena
A semantic-retrieval + code-editing MCP toolkit that exposes IDE-grade language-server operations (symbol lookup, references, refactors) to any MCP client. Works alongside Claude Code, Codex, JetBrains, and Cursor; Python-based, MIT-licensed, 1,632 forks, 104 open issues. Effectively "the IDE for your agent" — replacing grep + read-file scaffolding with proper LSP-aware tool calls. Best-installed MCP for giving agents semantic (not text-search) code understanding — measurable accuracy improvement on multi-file refactors.
context7
Upstash's MCP server that injects up-to-date, version-correct library documentation into LLMs and AI code editors — directly addressing the "model trained on stale docs" failure mode. It remains one of the most installed MCP servers in practice (it's wired into many editor setups, including this workspace) and keeps shipping under active development. The default answer to "my agent is using outdated API docs" — a staple MCP server that stays current.

AI Tools

Agent-Native
A framework for building agent-native applications where one `defineAction()` powers every surface: UI, agent, HTTP, MCP, A2A, and CLI. From BuilderIO, it's a "don't pick between apps or agents" approach that unifies app and agent entry points behind a single schema. A pragmatic answer to the app-versus-agent split — one action definition, six consumption surfaces.
Agent Governance Toolkit
Microsoft's public-preview toolkit for shipping autonomous agents to production: policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering, with coverage mapped to 10/10 of the OWASP Agentic Top 10. One of the few concrete, framework-level answers to agent safety and production governance.
lyogavin/airllm
Dramatically reduces inference memory so 70B LLMs run on a single 4GB GPU — without quantization, distillation, or pruning. The project claims 405B Llama 3.1 on 8GB, DeepSeek-V3 (671B) on ~12GB, and support for Kimi K3 (2.8T). Quickstart, config docs, macOS notes, and example notebooks are provided. Runs 70B-class open models on a 4GB GPU with no quantization — a practical answer to the GPU barrier.
addyosmani/agent-skills
Production-grade engineering skills for AI coding agents, encoding the workflows, quality gates, and best practices senior engineers use when building software. Skills are packaged so agents follow them consistently across DEFINE, PLAN, BUILD, VERIFY, REVIEW, and SHIP phases. A useful reference for teams standardizing how their agents build. Packages senior-engineer quality gates into agent skills — an easy win for engineering discipline.
BoundaryML/baml
A TypeScript-like programming language for agents where types persist at runtime and there's no `any` or unsafe casting — every feature is built to make agents make fewer mistakes. It compiles faster than Go and targets the messy boundary between structured program logic and LLM output. Typed, compile-time guarantees for agent I/O is a strong answer to the "unreliable JSON from the model" problem.
AI-Infra-Guard
A full-stack AI red-teaming platform from Tencent's Zhuque Lab that scans agents, skills, MCP servers, and AI infrastructure, and includes LLM jailbreak evaluation. It's one of the first tools to treat MCP servers and agent skills as first-class security surfaces. Security tooling that covers the whole AI stack — agents, skills, MCP, and infrastructure — at a time when that surface is expanding fast.
Ollama
The de-facto local model runner — a single CLI and API for getting up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma, and more on macOS, Windows, Linux, and Docker. It keeps trending because local inference remains central to agent and automation stacks. Still the fastest on-ramp to running open models locally for agents and automations.
Microsandbox
An easy, fast, local-first microVM runtime and library for running untrusted workloads — AI agents, user code, plugins, CI jobs, scrapers, and automation — with hardware-level isolation. Cross-platform on Linux and macOS, it gives agent developers a safe execution boundary without the overhead of full cloud sandboxes. Hardware-isolated local microVMs as a building block for running agent code and untrusted automation safely.
oMLX
oMLX runs local LLM inference on Apple Silicon Macs from a menu-bar app, using continuous batching and a two-tier cache — hot in memory, cold on SSD — so context from an earlier conversation stays reusable even after the model itself gets swapped out. Since going public in February it has picked up custom kernels for GLM-5.2, MiniMax M3, and Qwen3.5, and it's now experimenting with spreading inference across multiple Macs over Thunderbolt. Six months old and already past 20,000 stars — local Mac inference clearly still has unmet demand.
usestrix/strix
Open-source AI penetration testing tool with autonomous AI hackers that find and fix app vulnerabilities. Integrates seamlessly with GitHub Actions and CI/CD pipelines — automatically scans on every pull request and blocks insecure code before it reaches production. AI-powered security that runs in CI — shifts left on vulnerability detection with autonomous agents.
browser-use/video-use
Edit videos with Claude Code — 100% open source. Drop raw footage in a folder, chat with Claude Code, get final.mp4 back. Cuts filler words, dead air, and lets you edit talking heads, montages, tutorials, travel, and interviews without presets or menus. Turns a coding agent into a video editor — a novel application of agent capabilities to creative production.
kunchenguid/axi
A set of 10 design principles for building agent-ergonomic applications. Proposes AXI (Agent eXperience Interface) as a third paradigm alongside CLIs and MCP, claiming higher accuracy with lower token cost than both. A framework for designing tools that agents can use efficiently. Thought leadership on agent ergonomics — a design framework that could influence how agent-facing tools are built.
JuliusBrussee/caveman
A skill/plugin for Claude Code, Codex, Gemini CLI, and other coding agents that makes them "talk like a caveman" — cutting 65% of output tokens while maintaining answer quality. Multiple levels of compression available. Practical token optimization that directly reduces API costs for heavy agent users.
microsoft/flint-chart
A visualization intermediate language that lets AI agents create expressive, polished charts from simple, human-editable specs. Includes an MCP server guide. Designed to be the reliable charting layer for agent-generated visualizations. Microsoft's answer to the "make a chart" problem in AI agents — a focused, practical solution with MCP integration.
kyutai-labs/pocket-tts
A lightweight TTS application designed to run efficiently on CPUs — no GPU required. Supports Python 3.10–3.14 and is a single `pip install` away. Built by Kyutai, the non-profit AI research lab behind Moshi. CPU-only TTS that fits in any agent pipeline — removes the GPU barrier for voice output in local agents.
ECC
An agent-harness performance-optimization system layering skills, "instincts," memory, security, and research-first development on top of Claude Code, Codex, OpenCode, Cursor, and other harnesses. Created 2026-01-18, it has rocketed past 200k stars — the single highest-starred AI repo in this snapshot — riding the wave of developers trying to squeeze more reliability and structure out of their coding agents. The most-starred agent-tooling repo on GitHub right now — a harness-agnostic optimization layer the whole coding-agent crowd is piling into.
Odysseus
A self-hosted, local-first AI workspace — the DIY answer to the ChatGPT/Claude UI — bundling chat, an OpenCode-based agent, deep research, model "Cookbook" (hardware-aware model recommendation + serving via vLLM/llama.cpp/Ollama), persistent memory/skills (ChromaDB), plus email, calendar, notes, and a mobile PWA. Created 2026-05-31, it cleared ~18k stars almost overnight — the breakout new repo of this snapshot. The fastest-rising new project right now — a privacy-first, run-it-yourself workspace that replaces a stack of hosted AI tools.
ds4
A DeepSeek 4 Flash local inference engine for Metal and CUDA, from antirez (creator of Redis). Created 2026-05-06, it's a lean, readable C implementation in the spirit of his other from-scratch projects — and it's surged on both the model's release and the author's reputation for clear, hackable systems code. A high-signal pick for anyone running DeepSeek models locally. A from-scratch local inference engine for a hot model, written by one of open source's most respected systems engineers.
zerolang
"The programming language for agents," from Vercel Labs. Created 2026-05-15, it's an early, experimental take on a purpose-built language/runtime for expressing agent behavior — a notable bet that agents need their own primitives rather than being bolted onto general-purpose languages. Worth watching precisely because of who's behind it. A first-party Vercel Labs experiment in giving agents a native language — an early signal of where agent runtimes may go.
SkillOpt
A Microsoft research project that treats agent skills as something to *optimize*: a text-space optimizer that trains reusable natural-language skills for frozen LLM agents via trajectory-driven edits, validation-gated updates, and deployable `best_skill.md` artifacts. Created 2026-05-08, it's an early but credible take on self-improving skills without touching model weights. A first-party Microsoft approach to self-evolving agent skills — turning prompt/skill engineering into a measurable optimization loop.
firecrawl
An API to search, scrape, and crawl the web at scale and return LLM-ready markdown — the de facto web-access layer for agents and RAG pipelines. It handles JS-heavy sites, structured extraction, and full-site crawls behind a single call, which is why it shows up as a dependency in a growing share of agent stacks. Momentum remains high as agentic web tasks become a default workload. The cleanest "give my agent the web" primitive, and it keeps climbing.
colbymchenry/codegraph
Pre-indexed code knowledge graph for Claude Code, Codex, Gemini, Cursor, and Hermes Agent — ~16% cheaper, ~58% fewer tool calls, 100% local. A practical optimization that reduces token spend and tool calls across all major coding agents.
supermemoryai/supermemory
The memory and context layer for AI — #1 on LongMemEval, LoCoMo, and ConvoMem benchmarks. A state-of-the-art memory engine that can be used as a company or personal brain. Benchmark-leading AI memory infrastructure that solves the context persistence problem for agents.
LMCache/LMCache
A KV cache management layer for scalable LLM inference. Supercharges LLM serving with the fastest KV cache layer available. Recent updates include agentic workload benchmarks on AMD MI300X and a new multiprocess architecture release. KV cache optimization is critical for production LLM deployments, and LMCache is pushing the state of the art with real benchmarks.
huggingface/OpenEnv
An interface library for RL post-training with environments. Provides an end-to-end framework for creating, deploying, and using isolated execution environments for agentic RL training, built with Gymnasium-style APIs. Includes a featured example training LLMs to play BlackJack. From HuggingFace — a framework that bridges RL training and agentic environments, directly relevant to anyone doing agent post-training.
pydantic/monty
A minimal, secure Python interpreter written in Rust for use by AI. Avoids the cost, latency, and complexity of full container-based sandboxes for running AI-generated code. Still experimental but from the Pydantic team. A Rust-based sandboxed Python interpreter for AI code execution — addresses a core infrastructure need for agent tool use.
kenn-io/agentsview
Local-first session search, analytics, insights, and token use statistics for coding agents. Supports Claude Code, Codex, and 20+ other agents. One binary, no accounts, everything stays local. Fills a critical gap — observability for the multi-agent workflow era, with privacy-first local storage.
dmtrKovalenko/fff
The fastest and most accurate file search toolkit for AI agents, Neovim, Rust, C, Python, Bun, and NodeJS. Typo-resistant path and content search, frecency-ranked file access, background watcher, and lightweight in-memory content index. Claims to be way faster than ripgrep and fzf in long-running processes. File search is a bottleneck for agentic coding — this directly addresses agent speed and accuracy in large codebases.
google-labs-code/design.md
A format specification for describing a visual identity to coding agents. DESIGN.md gives agents a persistent, structured understanding of a design system using YAML frontmatter for design tokens and markdown for guidelines. Enables agents to maintain consistent visual output across sessions. Standardizes how agents understand design systems — a missing piece for production-grade agentic UI development.
jamiepine/voicebox
Open-source AI voice studio. Clone any voice, generate speech, dictate into any app, and talk to agents in voices you own. The full voice I/O stack runs locally on your machine. Includes CLI, API, and desktop app. A complete, local-first voice stack for agent interaction — relevant as voice agents gain traction.
every-app/open-seo
Open-source alternative to Semrush and Ahrefs. Pay-as-you-go SEO tool that you control. Connect with any agent like Claude Code, OpenClaw, or Hermes via pre-built skills. All-in-one SEO tool for you and your AI agent. An agent-native SEO tool that lets AI agents run SEO workflows directly — practical and timely.
mirage
A unified virtual filesystem for AI agents — a sandboxable layer that agents can read, write, and snapshot against without touching the real disk. Ships adapters for LangChain, the OpenAI Agents SDK, Claude Code, and shell agents, with a TypeScript core and Python clients. Created 2026-05-06 and cleared 2,200 stars within nine days, the largest single-week surge of any AI-tooling repo in the window. A clean "agent root filesystem" primitive plugs a gap every multi-tool agent harness has been hand-rolling.
how-to-train-your-gpt
A from-scratch PyTorch implementation of a modern GPT-style LLM where every line is commented and the architecture (attention, tokenisation, training loop) is walked through "explained like we are five." Created 2026-05-03 and picked up more than 1,500 stars in a week as a fresh entry in the "build-an-LLM" educational canon next to nanoGPT and llm.c. Cleanest new educational LLM build for engineers who want to understand the stack without wading through research code.
tokenspeed
A "speed-of-light" LLM inference engine in Python targeting Blackwell-class GPUs, with kernels tuned for gpt-oss, DeepSeek, Kimi, MiniMax, and Qwen out of the box. Created 2026-05-06 and crossed 1,000 stars in nine days as the open-source inference space keeps chasing vLLM and SGLang on raw throughput. Strongest new contender in the "fast inference for open weights" segment this week — worth watching against vLLM/SGLang benchmarks.
agent-skills-eval
A CLI test runner for agent skills in the agentskills.io style — feed it YAML/JSONL skill definitions plus eval cases and it scores them against any OpenAI-compatible endpoint. Created 2026-05-06; lands as the agentic-skills ecosystem matures past "marketplace of YAMLs" into "marketplace plus eval harness." Fills the eval gap for skill-based agents the same way `inspect_ai` filled it for general LLM evals.
html-anything
An agentic HTML editor bundling 75 skills × 9 surfaces (magazine, deck, poster, social-card, prototype, data report, Hyperframes) with sandboxed preview and one-click export to WeChat, X, Zhihu, HTML, and PNG. BYOK from Claude Code, Cursor, Codex, Gemini, Copilot, OpenCode, Qwen, or Aider — no first-party API key required. Created 2026-05-11 and cleared 1,900 stars in four days, the largest fresh launch in the AI-tooling tab this week. First credible "agent-writes-the-HTML, you ship it" editor with a serious skill library and a real sandbox — fills the gap between Claude Design knock-offs and freeform vibe-coding.
context-mode
A context-window optimiser for coding agents that sandboxes tool output and claims a 98% reduction across 15 platforms — Claude Code, Codex, Cursor, Copilot, OpenCode, Kiro, Antigravity, Zed, and more. Distributes as a plugin / hook layer rather than a wrapper, so existing harnesses keep their workflows. Created 2026-02-23 and still adding stars at multiple-hundred-per-week pace as agent operators discover token bills are now their #1 cost. Cost-per-task is the second axis the agent stack is now optimising, and `context-mode` is the most-installed open-source answer.
OpenMythos
A theoretical from-first-principles reconstruction of the "Claude Mythos" architecture in PyTorch / JAX, walking through attention, looped transformers, and training scaffolding using only the available research literature. Created 2026-04-18; passed 12,000 stars by 2026-05-15 on the strength of being the most readable open-source attempt at modelling Anthropic's published claims. Strongest current pick for engineers who want to understand the post-2025 frontier-LLM stack without waiting for a paper drop.
guizang-ppt-skill
An agent skill for generating polished HTML slide decks — editorial-magazine and Swiss layouts, image-prompt generation, social covers, and a WebGL / low-power presentation runtime — drop-in for Claude Code and Codex. Created 2026-04-23; nearly 9,000 stars and 733 forks at three weeks old as the consumer-facing "make me a deck" use case eats LangChain-flavoured pipelines. Best-in-class example of a single skill package outshipping a framework for a narrow, high-value workflow.
open-codesign
An open-source Claude Design alternative — Electron desktop app, prompt-to-prototype / slides / PDF flow, multi-model fanout (Claude, GPT, Gemini, Kimi, GLM, Ollama), BYOK, local-first, MIT. Imports your existing Claude Code or Codex API key with a click. Created 2026-04-18 and a near-mover of the "Claude Design clone" wave. Cleanest, most-installable answer to "I want Claude Design but I want to bring my own model and run it locally."
Claude-Code-Design-AI
A Claude Code plugin pitching itself as the AI UI/UX architect — screenshot-to-React, Figma-component generation, Tailwind / shadcn-ui scaffolding, wireframe rendering, SVG icon creation, dark-mode toggles, and responsive-layout tooling. Created 2026-05-13; cleared 370 stars in two days as the "Claude Design alternative" arms race kept compounding. Frames itself as the front-end-export bridge between Claude Design and a real React codebase — a missing piece in most of the clones.
journal-adapt-writing-skill
A Claude-Code skill that learns a target journal's writing conventions from its published papers, then rewrites a manuscript section-by-section to match. LaTeX-aware, economics-flavoured by default, MIT-licensed. Created 2026-05-13 and at 233 stars by 2026-05-15 — small in absolute terms, but a sharp instance of the "domain-specific skill" pattern eating prompt-engineering posts. Concrete proof that narrow, single-task skills outperform generic "academic writing" prompts when the workflow is well-defined.
forkd
A Rust-based microVM sandbox built on KVM that forks a warmed parent in 101 ms using copy-on-write snapshots — aimed squarely at agents that need a clean OS for every tool call. Created 2026-05-11. Sits in the same problem space as Firecracker / Mirage but trades ergonomics for raw fork latency. Lowest-latency open-source sandbox primitive for agents that have to spin up a clean kernel per task — paired with `mirage` (last week's #1) it forms a credible agent-runtime substrate.
Skill_Seekers
Converts documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection. Python, MIT-licensed, AST-parser-backed, OCR support for PDFs, web-scraping for docs sites, and a hosted companion at `skillseekersweb.com`. 1,402 forks and 97 open issues — community-driven, with daily pushes. Solves the "how do I get my private docs into a Claude skill" problem at scale — turns the skills marketplace from a hand-curated artifact into a generator pipeline.
agent-study
A 36-chapter agent-engineering curriculum (Chinese) — runnable Python files for ReAct loops, Claude Code reverse-engineering, MCP and A2A protocols, RAG, DSPy, and production observability — explicitly framed as interview prep. Created 2026-05-14 and at 264 stars / 20 forks four days later, with the interview-prep angle pushing it past the saturated "learn LangChain" cohort. Most pragmatic recent agent-engineering course material — fully runnable Python rather than slides, and ordered by what comes up in interviews now.
mempalace
Pitched as the best-benchmarked open-source AI memory system, and free. Created 2026-04-05, it's a Python + ChromaDB long-term memory layer for LLMs and agents that has already cleared 53k stars in under two months. Agent memory is the most contested infrastructure category of the quarter, and MemPalace is leading the open-source pack on benchmarks. Fastest-rising entry in the white-hot agent-memory category, with a benchmark-first pitch.
open-design
A local-first, open-source Claude Design alternative shipping as a native desktop app: 259+ skills, 142+ design systems, and sandboxed preview/export to HTML/PDF/PPTX/MP4 across web, desktop, mobile, slides, and video. Created 2026-04-28 and already near 56k stars, it rides the Figma-alternative + BYOK wave and plugs into 17+ agent CLIs. A serious open-source design surface for agents — local-first, multi-CLI, and moving fast.
hyperframes
"Write HTML. Render video. Built for agents." HeyGen's framework lets an agent emit HTML/GSAP and get back rendered video via Puppeteer + FFmpeg — a clean primitive for programmatic video generation. Created 2026-03-10, it's near 23k stars and is part of the broader "HyperFrames" pattern other tools (open-design, html-anything) are now adopting. Turns video into something an agent can author in HTML — a genuinely new output modality for LLM pipelines.
Crawl4AI
Crawl4AI turns web pages into clean markdown for RAG pipelines and agent context, handling dynamic content and structured extraction along the way. Version 0.9.3, released this cycle, is a security-only patch closing five coordinated-disclosure advisories — an arbitrary file write, an SSRF hole, a denial-of-service path in PDF processing, and two XSS bugs in the Docker Playground — with 33 additional bug fixes and no new features. a huge share of RAG and agent stacks route their web ingestion through this one project, so a coordinated security release here matters well beyond its own repo.
Khoj
Khoj is a self-hostable personal AI that chats with any local or cloud model, pulls answers from your own documents, and reaches you through a browser, Obsidian, Emacs, WhatsApp, or a desktop app. Its newest addition, Pipali, is an open-source AI coworker that runs on your own machine rather than a hosted server — a step past retrieval-and-chat toward something that does ongoing work. one of the few personal-AI projects that stays genuinely self-hostable while still expanding what it can do, instead of quietly becoming a funnel to a paid cloud tier.
Scientific Agent Skills
This is a library of 163 validated agent skills paired with access to over 100 scientific databases across biology, chemistry, medicine, and drug discovery, documented in an accompanying arXiv paper. It works with Cursor, Claude Code, Codex, and any agent that follows the open Agent Skills standard, plus a companion desktop co-scientist app that keeps data local. it's the difference between an agent that can summarize a paper and one that can actually run a multi-step research workflow against real databases.

AI Agents

ego-lite
A browser built for AI agent automation that shares your logged-in browser state with agents like Codex or Claude Code, letting them run multiple tasks in parallel "Spaces" while you work undisturbed. Zero cost, zero config. Directly solves the logged-in-session problem that blocks most agent browser automation today.
Speech To Speech
A low-latency, fully modular voice-agent pipeline — VAD → STT → LLM → TTS — exposed through an OpenAI Realtime-compatible WebSocket API. Every component is swappable, and the LLM slot accepts any OpenAI-compatible endpoint, from hosted providers to local models via HF Inference Providers. A maintained, local-first reference pipeline for building voice agents without vendor lock-in.
QwenLM/qwen-code
An open-source AI coding agent that lives in your terminal, agentic out of the box with Auto-Memory, Auto-Skills, SubAgents, Agent Teams, and built-in MCP support. The framework and Qwen models are fully open source, and dynamic workflows require zero setup. Docs ship in six languages. The open-source terminal coding agent with memory, skills, subagents, and MCP built in — trending as the default local agent pick.
Significant-Gravitas/AutoGPT
The open-source platform for AI agents: describe what you want done and AutoGPT builds the agent, runs it, and reports back. It positions self-hosting as a first-class path alongside hosted tiers. Still the most recognizable name in open agent frameworks, and it keeps trending. AutoGPT remains the broadest open path from a goal description to a finished agent run.
livekit/agents
Framework for building realtime, programmable voice agents that run server-side and can see, hear, and understand. A flexible integration ecosystem lets you mix and match speech, LLM, and transport components; the sibling AgentsJS project covers JS/TS. The default starting point for production voice-AI work. The reference framework for realtime voice agents that see and hear — core infrastructure for voice automation.
TencentCloud/TencentDB-Agent-Memory
A team-level memory hub for AI agents that turns conversations, docs, and code into four reusable assets: Chat Memory, Skill, LLM-Wiki, and Code-Graph. Memories are governed, shared, and equipped across agents and frameworks rather than locked inside one session. Includes benchmark results, technical implementation notes, and a roadmap. Memory is the missing layer for agent teams; this makes it structured, governed, and reusable.
anthropics/skills
Anthropic's public implementation of Agent Skills for Claude — folders of instructions, scripts, and resources that Claude loads dynamically to improve performance on specialized, repeatable tasks. This is the reference implementation behind the agentskills.io standard, so it's the baseline other skill packs and tools are being measured against. The canonical skills repo every agent builder will copy, extend, or target for compatibility.
google/skills
Agent Skills for Google products and technologies, including Google Cloud, installable via `npx skills add google/skills` with per-skill selection. It's under active development, but its existence confirms the skills format is going multi-vendor rather than staying Anthropic-specific. Signals that Agent Skills is becoming a cross-ecosystem standard, not a single-vendor format.
PrimeIntellect-ai/prime-agent
An open-source coding and research agent built around the Recursive Language Model (RLM) abstraction, which treats context as variables ("prompt-as-a-variable") to handle long-running autonomous tasks. Pairs with PRIME-RL and pi-mono for self-improvement loops, making it one of the more research-forward coding agents trending right now. A genuinely different architecture for long-horizon autonomy instead of another thin wrapper on a chat model.
paperclipai/paperclip
An open-source Node.js server and React UI for orchestrating a team of AI agents running business workflows — pitched as "if OpenClaw is an employee, Paperclip is the company." It's aimed at managing fleets of agents in one place rather than driving a single agent session. Early open-source take on the management layer that teams will need once multiple agents are doing real work.
elizaOS/eliza
An open-source TypeScript framework and product stack for autonomous AI agents — the monorepo includes the core runtime, the Eliza app, CLI, cloud services, native bridges, and first-party plugins. It's positioning itself as a bootable agentic operating system rather than a single-purpose agent. One of the most complete, actively developed agent frameworks, and it keeps showing up in trending for a reason.
Microsoft Agent Framework
An open, multi-language framework for building, orchestrating, and deploying production-grade AI agents and multi-agent workflows in Python and .NET. Microsoft built it for teams taking agents beyond prototypes into real deployments, with orchestration and workflow support baked in. It's the clearest first-party signal yet of where the agent framework landscape is consolidating. Cross-language, production-oriented multi-agent orchestration from Microsoft — a likely default for enterprise teams.
OpenViking
An open-source, self-evolving context database for AI agents from Volcengine that unifies agent memory, knowledge RAG, and skills in one store. It positions itself as the persistence and context layer for agents that need long-term memory and retrievable knowledge across sessions. Memory + RAG + skills in a single agent-facing database — a strong candidate for the agent context layer.
Apache Maka
Apache Maka (Incubating) is a local-first AI agent workspace that inspects projects and runs tools behind a sandbox boundary, recording model messages, tool calls, tool results, permission decisions, and termination events as an append-only log. The audit-first design makes agent behavior inspectable and reproducible. A sandboxed, local-first agent workspace with full audit logging — transparency by design.
ai-memory
ai-memory gives coding agents a memory that survives past the session and past the tool: quit Claude Code mid-task, open Codex in the same directory, and it picks up the architecture decisions, failed approaches, and open questions without you re-explaining any of it. It ships native binaries for macOS, Linux, and Windows via WSL2, and it's barely three months old — created in May, already past 4,600 stars. Solves the specific, annoying problem of losing context every time you switch coding agents mid-task.
ogulcancelik/herdr
An agent multiplexer that lives in your terminal — run all your coding agents (Claude Code, Codex, Cursor, etc.) in one terminal view. See who's blocked, working, or done at a glance. Agents run where they already run: your machine, a server, anywhere you can SSH. Solves the real pain of managing multiple parallel coding agents from a single terminal — a clear gap in the current agent workflow stack.
stablyai/orca
Orca is the Agent Development Environment (ADE) for working with a fleet of parallel agents. Run Codex, Claude Code, OpenCode, or Pi side-by-side — each in its own worktree, tracked in one place. Available on desktop and mobile with a mobile companion for monitoring and steering agents remotely. A dedicated desktop+mobile ADE for parallel agent fleets — fills the same gap as herdr but with a full GUI and mobile companion.
browser-use/browser-use
A library that makes websites accessible for AI agents, enabling reliable browser automation. The CLI 3.0 release gives coding agents a browser they can use. Supports Python with `uv add browser-use` or `pip install browser-use`, plus a cloud offering for stealth-enabled automation. The go-to open-source browser automation library for AI agents, now with a major CLI update.
alibaba/page-agent
A JavaScript in-page GUI agent that lets users control web interfaces with natural language. One script gives any web page its own AI agent — no browser extension, Python, or headless browser required. Everything happens in-page. Elegant approach to web agent integration — no infrastructure needed, just a script tag.
agentskills/agentskills
The specification and documentation for the Agent Skills open standard — a lightweight format for extending AI agent capabilities. A skill is simply a folder with a SKILL.md file containing metadata and instructions. This standard is being adopted by Google, Microsoft, and dozens of community projects. The emerging standard that multiple repos in this week's trending list are built upon — foundational infrastructure for the agent ecosystem.
TencentCloud/CubeSandbox
High-performance, secure sandbox service for AI agents built on RustVMM and KVM. Supports instant concurrent execution, single-node or multi-node deployment, and is designed specifically as a lightweight sandbox for running untrusted agent code safely. Addresses the critical need for safe code execution environments for AI agents, with production-grade performance.
LobeHub
Formerly LobeChat, now repositioned as a "Chief Agent Operator" — a surface for hiring, scheduling, and reporting on a whole team of agents running 7×24. It keeps a large knowledge-base, MCP, and multi-model (Claude, GPT, Gemini, DeepSeek) feature set, and the rename signals the broader shift from single-chat UIs to agent-fleet management. Active development continues on the `canary` branch. A mature, heavily-starred chat UI pivoting hard into agent-fleet orchestration — a bellwether for where AI front-ends are heading.
oh-my-openagent (omo)
Pitched as "the pickaxe for complex software engineering" — an agent harness purpose-built for large, gnarly codebases, with a TUI and broad model support (Claude, GPT, Gemini, OpenCode, Codex). Created 2025-12-03, it has climbed past 60k stars by targeting the exact pain point that generic agents struggle with: navigating and editing genuinely complex repos. A fast-rising harness aimed squarely at hard, real-world codebases rather than toy demos.
agents
A multi-harness agentic plugin marketplace spanning Claude Code, Codex CLI, Cursor, OpenCode, and Gemini CLI — subagents, skills, commands, and orchestration workflows in one installable collection. As teams standardize on reusable agent components, a cross-harness marketplace is exactly the kind of glue that keeps accumulating stars and contributions. The closest thing to a package manager for agent skills and subagents across every major coding harness.
hermes-agent
NousResearch's "agent that grows with you" — a long-running, model-agnostic agent harness that accumulates context and skills across sessions rather than starting cold each run. It sits alongside Claude Code, Codex, and OpenClaw as a coding/automation agent and is one of the fastest-accumulating agent repos on GitHub right now. The large open-issue count reflects an unusually active community shaping its direction. The reference open agent harness everyone is benchmarking against this quarter.
openclaude
Created on 2026-04-01, openclaude has accumulated 28k+ stars in roughly two months with the tagline "runs anywhere, uses anything" — a model-agnostic CLI agent that isn't locked to a single provider. Its fork count (8,600+) is unusually high relative to its age, suggesting heavy hacking and downstream experimentation. Worth watching as a lightweight, portable alternative to vendor-specific agent CLIs. Fastest-rising brand-new agent CLI in this snapshot.
goose
An open-source, extensible agent (Rust) that goes past code suggestions to actually install, execute, edit, and test against any LLM. It speaks both ACP and MCP, making it a flexible host for the broader tool ecosystem rather than a closed product. Steady commits and a clean extension model keep it a strong pick for teams wanting a self-hostable agent runtime. A mature, MCP-native agent runtime you can actually self-host and extend.
NVIDIA/OpenShell
A safe, private runtime for autonomous AI agents that provides sandboxed execution environments governed by declarative YAML policies. Prevents unauthorized file access, data exfiltration, and uncontrolled network activity. Addresses the critical unsolved problem of agent safety and sandboxing with an official NVIDIA solution.
NVIDIA/SkillSpector
Security scanner for AI agent skills. Detects vulnerabilities, malicious patterns, and security risks before installing agent skills. Research cited in the repo shows 26.1% of agent skills contain security issues — this tool addresses that gap directly for Claude Code, Codex CLI, Gemini CLI, and others. As the agent skills ecosystem explodes, a security scanner from NVIDIA is exactly what the community needs right now.
mvanhorn/last30days-skill
AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web, then synthesizes a grounded summary. An AI agent-led search engine scored by upvotes, likes, and real money — not editors. The v3 pipeline is documented in the README. A practical, multi-platform research skill that solves a real pain point for agents needing current, grounded information.
pbakaus/impeccable
Design guidance for AI coding agents. 1 skill, 23 commands, live browser iteration, and 41 deterministic detector rules for AI-generated frontend design. Install with `npx impeccable skills install`, then use `/impeccable init` inside your AI coding tool. Addresses the "AI generates ugly UIs" problem with deterministic design rules — practical for anyone shipping agent-built frontends.
microsoft/fara
Fara-7B: An efficient agentic model for computer use. A compact model designed for GUI automation and computer-use agent tasks, with a refreshed WebTailBench evaluation suite. Microsoft's bet on smaller, specialized agent models — relevant for anyone building computer-use agents without massive compute.
withastro/flue
The sandbox agent framework — a programmable TypeScript harness for building autonomous agents and AI workflows. Not another SDK; focuses on a harness pattern with skill imports and runtime routing. From the Astro team — brings their DX sensibility to agent frameworks, with a clean TypeScript-first approach.
Panniantong/Agent-Reach
Gives AI agents internet access via one CLI with zero API fees. Supports reading and searching Twitter, Reddit, YouTube, GitHub, Bilibili, and XiaoHongShu. Designed for agents that need real-time web data without managing individual API keys. Solves the "agent can't browse the web" problem with a unified, free CLI — a common gap in current agent toolchains.
NVIDIA/skills
Official, NVIDIA-verified skills for AI agents. Portable instruction sets that teach agents how to use NVIDIA software optimally, including CUDA-X libraries, AI Blueprints, and more. Part of NVIDIA's "capability governance" approach to agent skills. NVIDIA's verified skill library sets a quality bar for the agent skill marketplace.
awesome-agentic-ai-zh
A structured trilingual (Traditional Chinese / Simplified Chinese / English) learning roadmap for AI agents, organised as staged practicums with required readings, tagged with Claude Code, Claude Skills, MCP, and LLM-agent fundamentals. Created 2026-05-04, crossed 1,400 stars by 2026-05-15, and already has 151 forks — strong contributor velocity for a curriculum repo. The first "awesome-list" of the agentic-AI era to ship as a proper curriculum rather than an unranked URL dump.
Photo-agents
A computer-use agent stack built around "photographic" layered memory — the agent captures vision snapshots of the screen, writes its own skills from successful trajectories, and revises them over time. Created 2026-05-04 and shipped straight into the autonomous-agent / self-evolving-agent research thread, picking up over 850 stars by week 2. Concrete implementation of self-written skills tied to vision memory — a useful primitive for desktop and OS-control agents.
opensquilla
A token-efficient AI agent runtime pitching "same budget, higher intelligence density" — a programming model for skills + memory tuned around minimising tokens per useful action. Created 2026-05-06; one of several "openclaw"-family entrants this week pushing the cost-efficient agent angle and already past 800 stars. Worth tracking as cost-per-task becomes the next axis after raw capability in the agent runtime race.
Agent_Memory_Techniques
A 30-notebook cookbook covering conversation buffers, vector stores, knowledge graphs, episodic vs. semantic memory, MemGPT, Mem0, Letta, Zep, and Graphiti, with LoCoMo-benchmark code and production patterns. Created 2026-05-05 by the author of the popular `Prompt_Engineering` and `GenAI_Agents` cookbooks. Best single-source reference for the agent-memory toolchain that exploded over the last six months — directly runnable Jupyter, not just diagrams.
Auto-claude-code-research-in-sleep (ARIS)
A markdown-only autonomous ML-research skill pack — cross-model review loops, idea generation, experiment automation, and paper writing — that runs against Claude Code, Codex, OpenClaw, or any LLM agent with no framework lock-in. Created 2026-03-10 and now at 9.4k stars and 898 forks; the "no framework, just markdown skills" approach is what is driving the velocity. Cleanest expression of the "agents run while you sleep" pattern that's now common in academic ML workflows.
harmonist
A portable AI-agent orchestration framework with mechanical protocol enforcement — 186 agents, zero runtime dependencies, Python-only, designed to be dropped into Claude Code or Cursor without a sidecar. Created 2026-04-23 and crossed 1,600 stars / 342 forks by 2026-05-15, with the protocol-enforcement angle differentiating it from the LangGraph / CrewAI cohort. Strongest "single-binary multi-agent" pick this week — useful when you want an agent council without spinning up an event bus.
deer-flow
ByteDance's open-source long-horizon **SuperAgent harness** — sandboxes, memory, tools, skills, sub-agents, and a message gateway in one harness aimed at multi-minute-to-multi-hour tasks. LangChain / LangGraph-based with a polished Node.js / TypeScript control plane and a TikTok-grade product polish. Still shipping commits daily (819 open issues, 9,112 forks at this snapshot). The clearest answer right now to "if I want one open-source agent harness that a non-research team could actually deploy in production, what do I install?"
TradingAgents-astock
A multi-agent investment-research framework adapted for Chinese A-share markets — seven analyst agents (fundamentals, sentiment, technical, risk, etc.) running a bull / bear debate over A-share data sources (Dragon-Tiger list, hot-money flow, unlock schedules). A heavy fork-and-refit of the original `TradingAgents` repo. Created 2026-05-13; 96 forks in five days, with the China-domestic angle driving the velocity. Cleanest example of the "fork a generic agent framework, refit for a single regulated market" pattern that's eating bespoke quant pipelines.
hermes-agent-control-room
A "control room"-first template for operating Hermes agents — one VPS hosts an orchestrator that fan-outs to specialist teams and orchestrated workflows, with Shell-driven provisioning and a web console. Created 2026-05-15; 54 forks in three days. Targets the Nous Research / Hermes line specifically rather than being yet another LangGraph wrapper. Best-shaped current template for running Hermes as a *fleet* rather than as a single chat — the closest open-source thing to a Hermes ops console.
elephant-agent
A personal-model-first self-evolving AI agent. The headline pitch: the agent runs against your own local / fine-tuned model and rewrites its own context graph as it goes. Python, 23 open issues already (meaning real users on day three), and explicitly positioned against the "everyone runs Claude" default. Created 2026-05-15. Strongest fresh entry in the "BYO-model self-evolving agent" subgenre that has been quietly building behind the Claude-Code-skill wave.
autoresearch
Andrej Karpathy's experiment in fully autonomous ML research: AI agents that run, evaluate, and iterate on single-GPU nanochat training loops automatically. Created 2026-03-06, it rocketed past 84k stars on the strength of Karpathy's name plus a genuinely novel "agents doing the science" framing. It's a reference implementation more than a product, but it's the most-watched new AI-agent repo of the moment. The highest-signal new agent repo right now — autonomous research from the person who defined the genre.
nanobot
From HKUDS (the lab behind several popular RAG/agent projects), nanobot is a lightweight open-source AI agent for tools, chats, and workflows. Created 2026-02-01, it has crossed 43k stars by being the "small, hackable, model-agnostic" option in a field of heavyweight harnesses — easy to read, easy to extend. The credible lightweight agent for people who want to understand and modify their harness, not just run it.
ruflo
ruvnet's agent-orchestration platform for Claude: multi-agent swarms, autonomous workflow coordination, self-learning swarm intelligence, RAG, and native Claude Code / Codex integration. It's the most prominent of the "swarm" orchestration frameworks and continues to draw heavy engagement and contribution activity. The go-to open-source orchestration layer for multi-agent Claude swarms.
AgentMemory
AgentMemory records what a coding agent did during a session, compresses it, and feeds the relevant parts back into the next one — across Claude Code, Copilot CLI, Cursor, Codex, and most other MCP clients. It combines BM25 keyword search, vector embeddings, and a knowledge graph rather than relying on any single retrieval method. persistent memory that survives a restart, instead of another status file the agent forgets to read.

Integrations

OmniRoute
A free MIT-licensed AI gateway exposing one endpoint to 290+ providers (90+ free) and 500+ models, with clients for Claude Code, Codex, Cursor, OpenCode, Cline, and Copilot. Includes quota-aware auto-fallback, RTK+Caveman compression that claims 15–95% token savings, and MCP/A2A support. Drop-in multi-provider routing with fallback and token compression — practical cost infrastructure for agent CLIs.
NVIDIA-NeMo/Switchyard
A Rust proxy and library for LLM traffic that routes requests across providers while preserving native OpenAI and Anthropic API compatibility. It also translates between the two API formats, records operational metrics, and supports benchmarking plus cost/performance optimization — useful as a drop-in layer for apps that want provider flexibility without rewiring their SDK calls. Solves the practical "which model/provider should this call hit" problem with a fast, observability-aware proxy.
Claude Plugins — Community
Anthropic's read-only mirror of the community plugin marketplace for Claude Cowork and Claude Code; the repo's `marketplace.json` is the canonical list of community-contributed plugins. It's the discovery surface for extending Claude's agent tooling, with submissions routed through the official plugin directory. The official entry point to the Claude plugin ecosystem — worth watching for any Claude Code/Cowork developer.
openai/codex-plugin-cc
Use Codex from inside Claude Code for code reviews or to delegate tasks. Provides `/codex:review` for read-only reviews and `/codex:adversarial-review` for steerable challenge-based reviews. Bridges OpenAI's Codex into the Claude Code workflow. Cross-platform agent integration from OpenAI itself — lets Claude Code users leverage Codex without leaving their workflow.
claude-mem
Persistent memory across sessions for coding agents: it captures what your agent did, compresses it with the model, and re-injects the relevant slice into future sessions. It integrates broadly — Claude Code, OpenClaw, Codex, Gemini, Copilot, OpenCode — using ChromaDB/SQLite under the hood. Memory is one of the hottest agent problems right now, and this is among the most-starred open implementations. Drop-in long-term memory that works across nearly every coding agent.
agentgateway/agentgateway
The first complete connectivity solution for agentic AI — an open source proxy built on MCP and A2A protocols providing drop-in security, observability, and governance for agent-to-LLM, agent-to-tool, and agent-to-agent communication. Solves the emerging need for a universal proxy layer across all agent frameworks and protocols.
openpets
An Electron + Bun desktop pet that connects to Claude Code (and other coding agents) via MCP and animates its mood based on the agent's live status — running tool, waiting on user, errored, idle. Created 2026-05-04 with a companion repo `alvinunreal/claude-pets` (now archived) that handles the Claude Code hook side. Best example this week of the "ambient awareness for long-running coding agents" UX problem getting a playful native solution.
FreeLLMAPI
FreeLLMAPI aggregates free-tier access from 34 LLM providers behind a single OpenAI-compatible endpoint, covering 635 model endpoints and roughly 7.4 billion tokens a month. A router picks a live model per request and fails over automatically when one provider hits its cap, and the model catalog updates itself from a signed feed instead of requiring a fresh install. it solves the tedious part of running on free tiers — tracking which of thirty-some providers still has quota left — instead of just being one more provider.

Automations

T3 Code
An open "agent harness control surface" that lets you control the agents on your machine — Claude Code, Codex, Cursor, Grok Build, and OpenCode — from iOS/Android, a web app, or an Electron desktop app. It works with the subscriptions and agents you already have set up on your computer. Cross-harness agent orchestration from your phone or desktop without switching coding agents.
OpenSRE
An open-source framework for building AI SRE agents, including the training and evaluation environment they need to improve. Connects 60+ tools you already run, lets you define custom workflows, and investigates incidents on your own infrastructure — currently in public alpha. A rare open take on agentic incident response and reliability workflows, not just code generation.
Scrapling
An adaptive Python web-scraping framework that scales from a single request to a full crawl, with stealth/anti-bot handling, Playwright integration, and a built-in MCP server so agents can scrape directly. As agent-driven data collection becomes routine, a scraper that ships its own MCP interface is hitting at the right time. Production-grade scraping with native MCP — the data-ingestion layer for web-aware agents.

Security

Anthropic Cybersecurity Skills
Despite the name, this is an independent community project, not something Anthropic built or endorses. It packages 817 structured playbooks that walk an AI coding agent through a specific security task — a phishing simulation, an ATT&CK-mapped incident response, an AI red-team check — with each one tied back to a recognized framework like MITRE ATT&CK, NIST CSF 2.0, or MITRE ATLAS, so an agent follows a documented procedure instead of improvising. It added 94 new skills mapped to MITRE's Fight Fraud framework this spring and now spans 29-plus security domains across 20-plus agent platforms. Last covered here 2026-06-29 as an honorable mention; it earns a full return on that framework expansion and the jump in adoption. The largest structured security-skill library built specifically for AI agents, and it's still growing.
Shannon
Shannon is an autonomous pentester for web apps and APIs: it reads your source code, maps out attack paths, and runs real exploits against your own application to prove a vulnerability exists before it reaches production. Version 3.0, released this cycle, adds deeper source analysis, a rebuilt CLI, native CI/CD hooks, and SARIF and PDF report output aimed at security teams rather than a single terminal session. it turns "this endpoint looks risky" into a verified exploit chain, which is a different and more useful thing than a static scan.

AI/ML

Modular Platform
Modular's monorepo bundles the Mojo language and compiler with MAX, an inference server and model-pipeline framework meant to run AI workloads across CPUs and GPUs without locking developers into one vendor's hardware. It's one of the more serious attempts at a from-scratch AI systems stack rather than a wrapper around existing runtimes, with an OpenAI-compatible serving layer and a standard library the company is opening to outside contributors piece by piece. A ground-up AI systems stack, still expanding what parts are open for outside contribution.
Semantica
Semantica builds a knowledge graph out of enterprise data and keeps a record of how every conclusion was reached, aimed at industries like lending where an agent's decision has to survive a regulator asking why. The latest release, v0.6.6, added SSRF protections, a first-class CrewAI integration, and a GDPR-style purge that can retract a record along with everything derived from it. Decision provenance for agents in regulated industries — an audit trail as a first-class feature, not an afterthought.
FastVideo
FastVideo, out of Hao AI Lab, is a training and inference framework built to make video-generation models faster to run and cheaper to fine-tune, packaging distilled checkpoints that trade a sliver of quality for a large speedup. The newest of those, FastH3, landed three days before this list — a 4-step distilled MiniMax-H3 checkpoint that generates synchronized video and audio together instead of as separate passes. One of the few projects treating video-generation speed as its own research problem, with checkpoints to show for it.
KTransformers
KTransformers optimizes LLM inference and fine-tuning across mixed CPU-GPU hardware, aiming to let a single consumer GPU run models that would otherwise need a rack. On August 26, 2026 the project added native GLM-5.3-flash support with a 1M-token context window on consumer GPUs, followed within days by AVX512 CPU-only LoRA fine-tuning for x86 servers that lack AMX. most inference-optimization projects announce support for a new model months after release — this one shipped it within days.
WebLLM
WebLLM runs LLM inference directly inside a browser tab using WebGPU, with no server call once the model is loaded, and it mirrors the OpenAI API closely enough that existing client code needs little rewriting. It supports streaming, JSON mode, and a growing list of open model families. it's still one of the only credible paths to a private, fully offline AI feature that ships as an ordinary web page.

Data

Data Formulator
Data Formulator, out of Microsoft Research, splits the work of building a chart between the analyst and the model: point it at a spreadsheet, a database, or a screenshot, ask a question in plain language, and it handles the data wrangling while you steer the visualization. Its "Data Threads" feature lets you branch off a follow-up question without losing the path that got you there, which is the part most chat-based analysis tools get wrong. Treats exploratory data analysis as a branching conversation instead of one linear chat log.
turbovec
turbovec is a Rust vector index with Python bindings, built on Google Research's TurboQuant quantizer. The pitch is blunt: a 10-million-document corpus that takes 31 GB as float32 fits in 4 GB here, and it still searches faster than FAISS in the author's benchmarks, with no training phase and no rebuild required as the corpus grows. It recently picked up LangChain, LlamaIndex, and Haystack integrations, plus hand-written SIMD kernels for both ARM and x86. A FAISS alternative that skips the training step entirely and still wins on speed and memory.
LanceDB
LanceDB is an embedded, multimodal vector database built on the Lance columnar format, so one library handles text, image, video, and point-cloud embeddings alongside SQL and full-text search rather than bolting a vector index onto a relational store. It has cut regular releases for years — the latest, v0.37.1, landed August 10 — and stays close to LangChain and LlamaIndex as those frameworks evolve. One of the more mature, actively maintained multimodal vector stores, still shipping monthly.

Developer Tools

GitNexus
GitNexus indexes a codebase — GitHub, GitLab, Azure repos, or a local ZIP — entirely in the browser, building a knowledge graph of every dependency, call chain, and cluster without sending code to a server. It exposes that graph through MCP tools so an agent can trace impact and blast radius across files before touching anything, rather than inferring structure from a directory listing. it gives agents something closer to an actual mental model of a codebase, not just more text to search through.
Cursor Plugins
This is Cursor's own plugin specification plus its first batch of official plugins, each one a standalone directory with a `.cursor-plugin/plugin.json` manifest. The initial set covers things like automated branch review, incremental AGENTS.md memory updates, and a scaffolder for building new plugins. Cursor is formalizing an extension surface instead of leaving the pattern to whatever the community improvises.