← Top AI & Dev Repos
Bookmarks 2026-08-25

Top 10 GitHub Repos — 2026-08-25

Data infrastructure for AI is the throughline this week: Data Formulator, turbovec, and LanceDB each tackle a different piece of the same problem, turning oversized or messy data into something a…

Top 10 GitHub Repos — 2026-08-25
Open report
Add to your browser
Import these repos into your browser (15 KB) Import the entire archive (197 KB) How to import
  1. Click the button to download the -bookmarks.html file.
  2. Chrome / Edge / Brave: Bookmarks → Import bookmarks and settings → Favorites or bookmarks HTML file → pick the downloaded file.
  3. Firefox: Bookmarks → Manage Bookmarks → Import and Backup → Import Bookmarks from HTML.
  4. Safari: File → Import From → Bookmarks HTML File.
They import as one tidy folder (with category subfolders) you can delete anytime.

Top 10 GitHub Repos — 2026-08-25

As of: 2026-08-25 Focus: AI tools, automations, agents, integrations, MCP, CLIs, and developer tooling. Snapshot: Current trending picks at the moment of generation — not a historical recap.


Overview

Data infrastructure for AI is the throughline this week: Data Formulator, turbovec, and LanceDB each tackle a different piece of the same problem, turning oversized or messy data into something a model or an analyst can actually search. Local inference keeps pushing further to the edge too, with oMLX now experimenting with splitting inference across multiple Macs over Thunderbolt. Agent tooling is less about new frameworks this week and more about memory — ai-memory and Entire CLI both try to preserve a record of what an agent did, from opposite directions (session handoff versus git history). Anthropic Cybersecurity Skills returns after eight weeks off this list, back on the strength of a fraud-framework expansion and a jump past 31,000 stars. The javascript.xml trending feed failed to parse this week, so this list draws from five of the usual six category feeds.


1. Anthropic Cybersecurity Skills

Despite the name, this is an independent community project, not something Anthropic built or endorses. It packages 817 structured playbooks that walk an AI coding agent through a specific security task — a phishing simulation, an ATT&CK-mapped incident response, an AI red-team check — with each one tied back to a recognized framework like MITRE ATT&CK, NIST CSF 2.0, or MITRE ATLAS, so an agent follows a documented procedure instead of improvising. It added 94 new skills mapped to MITRE's Fight Fraud framework this spring and now spans 29-plus security domains across 20-plus agent platforms. Last covered here 2026-06-29 as an honorable mention; it earns a full return on that framework expansion and the jump in adoption.

Why it's on the list: The largest structured security-skill library built specifically for AI agents, and it's still growing.


2. Modular Platform

Modular's monorepo bundles the Mojo language and compiler with MAX, an inference server and model-pipeline framework meant to run AI workloads across CPUs and GPUs without locking developers into one vendor's hardware. It's one of the more serious attempts at a from-scratch AI systems stack rather than a wrapper around existing runtimes, with an OpenAI-compatible serving layer and a standard library the company is opening to outside contributors piece by piece.

Why it's on the list: A ground-up AI systems stack, still expanding what parts are open for outside contribution.


3. oMLX

oMLX runs local LLM inference on Apple Silicon Macs from a menu-bar app, using continuous batching and a two-tier cache — hot in memory, cold on SSD — so context from an earlier conversation stays reusable even after the model itself gets swapped out. Since going public in February it has picked up custom kernels for GLM-5.2, MiniMax M3, and Qwen3.5, and it's now experimenting with spreading inference across multiple Macs over Thunderbolt.

Why it's on the list: Six months old and already past 20,000 stars — local Mac inference clearly still has unmet demand.


4. Data Formulator

Data Formulator, out of Microsoft Research, splits the work of building a chart between the analyst and the model: point it at a spreadsheet, a database, or a screenshot, ask a question in plain language, and it handles the data wrangling while you steer the visualization. Its "Data Threads" feature lets you branch off a follow-up question without losing the path that got you there, which is the part most chat-based analysis tools get wrong.

Why it's on the list: Treats exploratory data analysis as a branching conversation instead of one linear chat log.


5. turbovec

turbovec is a Rust vector index with Python bindings, built on Google Research's TurboQuant quantizer. The pitch is blunt: a 10-million-document corpus that takes 31 GB as float32 fits in 4 GB here, and it still searches faster than FAISS in the author's benchmarks, with no training phase and no rebuild required as the corpus grows. It recently picked up LangChain, LlamaIndex, and Haystack integrations, plus hand-written SIMD kernels for both ARM and x86.

Why it's on the list: A FAISS alternative that skips the training step entirely and still wins on speed and memory.


6. LanceDB

LanceDB is an embedded, multimodal vector database built on the Lance columnar format, so one library handles text, image, video, and point-cloud embeddings alongside SQL and full-text search rather than bolting a vector index onto a relational store. It has cut regular releases for years — the latest, v0.37.1, landed August 10 — and stays close to LangChain and LlamaIndex as those frameworks evolve.

Why it's on the list: One of the more mature, actively maintained multimodal vector stores, still shipping monthly.


7. Semantica

Semantica builds a knowledge graph out of enterprise data and keeps a record of how every conclusion was reached, aimed at industries like lending where an agent's decision has to survive a regulator asking why. The latest release, v0.6.6, added SSRF protections, a first-class CrewAI integration, and a GDPR-style purge that can retract a record along with everything derived from it.

Why it's on the list: Decision provenance for agents in regulated industries — an audit trail as a first-class feature, not an afterthought.


8. Entire CLI

Entire CLI hooks into git and records what an AI coding agent actually did during a session — the prompts, the files touched, the tool calls — then indexes that transcript alongside the commit it produced. You can resume a session from any earlier checkpoint, or hand a teammate the full "why did the code change" story instead of just a diff. It shipped v0.10.2 on August 19, six days before this list went out, and has already logged nearly 8,000 commits of its own.

Why it's on the list: Turns agent sessions into a searchable part of git history instead of a conversation that evaporates.


9. ai-memory

ai-memory gives coding agents a memory that survives past the session and past the tool: quit Claude Code mid-task, open Codex in the same directory, and it picks up the architecture decisions, failed approaches, and open questions without you re-explaining any of it. It ships native binaries for macOS, Linux, and Windows via WSL2, and it's barely three months old — created in May, already past 4,600 stars.

Why it's on the list: Solves the specific, annoying problem of losing context every time you switch coding agents mid-task.


10. FastVideo

FastVideo, out of Hao AI Lab, is a training and inference framework built to make video-generation models faster to run and cheaper to fine-tune, packaging distilled checkpoints that trade a sliver of quality for a large speedup. The newest of those, FastH3, landed three days before this list — a 4-step distilled MiniMax-H3 checkpoint that generates synchronized video and audio together instead of as separate passes.

Why it's on the list: One of the few projects treating video-generation speed as its own research problem, with checkpoints to show for it.


Honorable Mentions

  • ZSeven-W/openpencil — AI-native vector design tool where concurrent agent teams build UI directly on the canvas from a prompt.
  • compozy/compozy — Coordinates coding-agent CLIs you already run (Claude Code, Codex, Gemini CLI, Cursor) into a team that splits and hands off work.
  • eneskirca/nodeterm — Node-based terminal manager that turns parallel AI agent sessions into draggable nodes on an infinite canvas.
  • marin-community/marin — Stanford-backed open research platform training a large mixture-of-experts model in the open, checkpoints and all.
  • future-agi/future-agi — Open-source tracing, eval, and guardrail platform for catching agent hallucinations before they ship; nightly builds, not yet stable.

Sources

More from Bookmarks