Top 10 GitHub Repos — 2026-09-04
Nothing on this list is a chatbot wrapper. The set skews toward infrastructure that other agents depend on: Crawl4AI shipped a security release patching five coordinated-disclosure advisories…
- Click the button to download the
-bookmarks.htmlfile. - Chrome / Edge / Brave: Bookmarks → Import bookmarks and settings → Favorites or bookmarks HTML file → pick the downloaded file.
- Firefox: Bookmarks → Manage Bookmarks → Import and Backup → Import Bookmarks from HTML.
- Safari: File → Import From → Bookmarks HTML File.
Top 10 GitHub Repos — 2026-09-04
As of: 2026-09-04 Focus: AI tools, automations, agents, integrations, MCP, CLIs, and developer tooling. Snapshot: Current trending picks at the moment of generation — not a historical recap.
Overview
Nothing on this list is a chatbot wrapper. The set skews toward infrastructure that other agents depend on: Crawl4AI shipped a security release patching five coordinated-disclosure advisories, KTransformers added native support for a 1M-token model days after it launched, and GitNexus builds a structural knowledge graph so agents stop guessing at blast radius from a flat file tree. Three entries are agent-facing layers rather than end-user products — a scientific skills library, a cross-agent memory system, and a free-tier LLM router — which says something about where the effort is going right now: less on the interface, more on what the agent can see and remember. Security tooling also gets real airtime this week with Shannon's 3.0 release adding CI/CD integration and SARIF output, aimed at teams rather than solo terminal users.
1. Crawl4AI
- Repo: https://github.com/unclecode/crawl4ai
- Stars: 81,300
- Category: AI Tools
- Status: Popular & active
Crawl4AI turns web pages into clean markdown for RAG pipelines and agent context, handling dynamic content and structured extraction along the way. Version 0.9.3, released this cycle, is a security-only patch closing five coordinated-disclosure advisories — an arbitrary file write, an SSRF hole, a denial-of-service path in PDF processing, and two XSS bugs in the Docker Playground — with 33 additional bug fixes and no new features.
Why it's on the list: a huge share of RAG and agent stacks route their web ingestion through this one project, so a coordinated security release here matters well beyond its own repo.
2. KTransformers
- Repo: https://github.com/kvcache-ai/ktransformers
- Stars: 19,500
- Category: AI/ML
- Status: Trending
KTransformers optimizes LLM inference and fine-tuning across mixed CPU-GPU hardware, aiming to let a single consumer GPU run models that would otherwise need a rack. On August 26, 2026 the project added native GLM-5.3-flash support with a 1M-token context window on consumer GPUs, followed within days by AVX512 CPU-only LoRA fine-tuning for x86 servers that lack AMX.
Why it's on the list: most inference-optimization projects announce support for a new model months after release — this one shipped it within days.
3. Khoj
- Repo: https://github.com/khoj-ai/khoj
- Stars: 37,100
- Category: AI Tools
- Status: Trending
Khoj is a self-hostable personal AI that chats with any local or cloud model, pulls answers from your own documents, and reaches you through a browser, Obsidian, Emacs, WhatsApp, or a desktop app. Its newest addition, Pipali, is an open-source AI coworker that runs on your own machine rather than a hosted server — a step past retrieval-and-chat toward something that does ongoing work.
Why it's on the list: one of the few personal-AI projects that stays genuinely self-hostable while still expanding what it can do, instead of quietly becoming a funnel to a paid cloud tier.
4. GitNexus
- Repo: https://github.com/abhigyanpatwari/GitNexus
- Stars: 47,000
- Category: Developer Tools
- Status: Trending
GitNexus indexes a codebase — GitHub, GitLab, Azure repos, or a local ZIP — entirely in the browser, building a knowledge graph of every dependency, call chain, and cluster without sending code to a server. It exposes that graph through MCP tools so an agent can trace impact and blast radius across files before touching anything, rather than inferring structure from a directory listing.
Why it's on the list: it gives agents something closer to an actual mental model of a codebase, not just more text to search through.
5. WebLLM
- Repo: https://github.com/mlc-ai/web-llm
- Stars: 19,000
- Category: AI/ML
- Status: Popular & active
WebLLM runs LLM inference directly inside a browser tab using WebGPU, with no server call once the model is loaded, and it mirrors the OpenAI API closely enough that existing client code needs little rewriting. It supports streaming, JSON mode, and a growing list of open model families.
Why it's on the list: it's still one of the only credible paths to a private, fully offline AI feature that ships as an ordinary web page.
6. Cursor Plugins
- Repo: https://github.com/cursor/plugins
- Stars: 6,800
- Category: Developer Tools
- Status: New
This is Cursor's own plugin specification plus its first batch of official plugins, each one a standalone directory with a .cursor-plugin/plugin.json manifest. The initial set covers things like automated branch review, incremental AGENTS.md memory updates, and a scaffolder for building new plugins.
Why it's on the list: Cursor is formalizing an extension surface instead of leaving the pattern to whatever the community improvises.
7. Scientific Agent Skills
- Repo: https://github.com/K-Dense-AI/scientific-agent-skills
- Stars: 42,500
- Category: AI Tools
- Status: New
This is a library of 163 validated agent skills paired with access to over 100 scientific databases across biology, chemistry, medicine, and drug discovery, documented in an accompanying arXiv paper. It works with Cursor, Claude Code, Codex, and any agent that follows the open Agent Skills standard, plus a companion desktop co-scientist app that keeps data local.
Why it's on the list: it's the difference between an agent that can summarize a paper and one that can actually run a multi-step research workflow against real databases.
8. AgentMemory
- Repo: https://github.com/rohitg00/agentmemory
- Stars: 28,000
- Category: AI Agents
- Status: New
AgentMemory records what a coding agent did during a session, compresses it, and feeds the relevant parts back into the next one — across Claude Code, Copilot CLI, Cursor, Codex, and most other MCP clients. It combines BM25 keyword search, vector embeddings, and a knowledge graph rather than relying on any single retrieval method.
Why it's on the list: persistent memory that survives a restart, instead of another status file the agent forgets to read.
9. FreeLLMAPI
- Repo: https://github.com/tashfeenahmed/freellmapi
- Stars: 24,200
- Category: Integrations
- Status: New
FreeLLMAPI aggregates free-tier access from 34 LLM providers behind a single OpenAI-compatible endpoint, covering 635 model endpoints and roughly 7.4 billion tokens a month. A router picks a live model per request and fails over automatically when one provider hits its cap, and the model catalog updates itself from a signed feed instead of requiring a fresh install.
Why it's on the list: it solves the tedious part of running on free tiers — tracking which of thirty-some providers still has quota left — instead of just being one more provider.
10. Shannon
- Repo: https://github.com/KeygraphHQ/shannon
- Stars: 47,700
- Category: Security
- Status: New
Shannon is an autonomous pentester for web apps and APIs: it reads your source code, maps out attack paths, and runs real exploits against your own application to prove a vulnerability exists before it reaches production. Version 3.0, released this cycle, adds deeper source analysis, a rebuilt CLI, native CI/CD hooks, and SARIF and PDF report output aimed at security teams rather than a single terminal session.
Why it's on the list: it turns "this endpoint looks risky" into a verified exploit chain, which is a different and more useful thing than a static scan.
Honorable Mentions
- JetBrains/go-modern-guidelines — JetBrains-authored guidelines that stop coding agents from writing Go the way it was written five versions ago.
- DietrichGebert/ponytail — a skill that pushes an agent toward the smallest correct diff instead of the most elaborate one, with benchmarks to back the claim.
- nashsu/llm_wiki — a desktop app that builds a standing wiki from your documents once instead of re-running retrieval on every question.
- MadsLorentzen/ai-job-search — a fork-it job-search framework built on Claude Code slash commands, written by a geophysicist who used it on his own search.
Sources
- Crawl4AI — https://github.com/unclecode/crawl4ai
- KTransformers — https://github.com/kvcache-ai/ktransformers
- Khoj — https://github.com/khoj-ai/khoj
- GitNexus — https://github.com/abhigyanpatwari/GitNexus
- WebLLM — https://github.com/mlc-ai/web-llm
- Cursor Plugins — https://github.com/cursor/plugins
- Scientific Agent Skills — https://github.com/K-Dense-AI/scientific-agent-skills
- AgentMemory — https://github.com/rohitg00/agentmemory
- FreeLLMAPI — https://github.com/tashfeenahmed/freellmapi
- Shannon — https://github.com/KeygraphHQ/shannon
- JetBrains Go Modern Guidelines — https://github.com/JetBrains/go-modern-guidelines
- Ponytail — https://github.com/DietrichGebert/ponytail
- LLM Wiki — https://github.com/nashsu/llm_wiki
- AI Job Search — https://github.com/MadsLorentzen/ai-job-search
More from Bookmarks