Bookmarks
Top 10 GitHub Repos — August 25, 2026
Security
- Anthropic Cybersecurity Skills
- Despite the name, this is an independent community project, not something Anthropic built or endorses. It packages 817 structured playbooks that walk an AI coding agent through a specific security task — a phishing simulation, an ATT&CK-mapped incident response, an AI red-team check — with each one tied back to a recognized framework like MITRE ATT&CK, NIST CSF 2.0, or MITRE ATLAS, so an agent follows a documented procedure instead of improvising. It added 94 new skills mapped to MITRE's Fight Fraud framework this spring and now spans 29-plus security domains across 20-plus agent platforms. Last covered here 2026-06-29 as an honorable mention; it earns a full return on that framework expansion and the jump in adoption. The largest structured security-skill library built specifically for AI agents, and it's still growing.
AI/ML
- Modular Platform
- Modular's monorepo bundles the Mojo language and compiler with MAX, an inference server and model-pipeline framework meant to run AI workloads across CPUs and GPUs without locking developers into one vendor's hardware. It's one of the more serious attempts at a from-scratch AI systems stack rather than a wrapper around existing runtimes, with an OpenAI-compatible serving layer and a standard library the company is opening to outside contributors piece by piece. A ground-up AI systems stack, still expanding what parts are open for outside contribution.
- Semantica
- Semantica builds a knowledge graph out of enterprise data and keeps a record of how every conclusion was reached, aimed at industries like lending where an agent's decision has to survive a regulator asking why. The latest release, v0.6.6, added SSRF protections, a first-class CrewAI integration, and a GDPR-style purge that can retract a record along with everything derived from it. Decision provenance for agents in regulated industries — an audit trail as a first-class feature, not an afterthought.
- FastVideo
- FastVideo, out of Hao AI Lab, is a training and inference framework built to make video-generation models faster to run and cheaper to fine-tune, packaging distilled checkpoints that trade a sliver of quality for a large speedup. The newest of those, FastH3, landed three days before this list — a 4-step distilled MiniMax-H3 checkpoint that generates synchronized video and audio together instead of as separate passes. One of the few projects treating video-generation speed as its own research problem, with checkpoints to show for it.
AI Tools
- oMLX
- oMLX runs local LLM inference on Apple Silicon Macs from a menu-bar app, using continuous batching and a two-tier cache — hot in memory, cold on SSD — so context from an earlier conversation stays reusable even after the model itself gets swapped out. Since going public in February it has picked up custom kernels for GLM-5.2, MiniMax M3, and Qwen3.5, and it's now experimenting with spreading inference across multiple Macs over Thunderbolt. Six months old and already past 20,000 stars — local Mac inference clearly still has unmet demand.
Data
- Data Formulator
- Data Formulator, out of Microsoft Research, splits the work of building a chart between the analyst and the model: point it at a spreadsheet, a database, or a screenshot, ask a question in plain language, and it handles the data wrangling while you steer the visualization. Its "Data Threads" feature lets you branch off a follow-up question without losing the path that got you there, which is the part most chat-based analysis tools get wrong. Treats exploratory data analysis as a branching conversation instead of one linear chat log.
- turbovec
- turbovec is a Rust vector index with Python bindings, built on Google Research's TurboQuant quantizer. The pitch is blunt: a 10-million-document corpus that takes 31 GB as float32 fits in 4 GB here, and it still searches faster than FAISS in the author's benchmarks, with no training phase and no rebuild required as the corpus grows. It recently picked up LangChain, LlamaIndex, and Haystack integrations, plus hand-written SIMD kernels for both ARM and x86. A FAISS alternative that skips the training step entirely and still wins on speed and memory.
- LanceDB
- LanceDB is an embedded, multimodal vector database built on the Lance columnar format, so one library handles text, image, video, and point-cloud embeddings alongside SQL and full-text search rather than bolting a vector index onto a relational store. It has cut regular releases for years — the latest, v0.37.1, landed August 10 — and stays close to LangChain and LlamaIndex as those frameworks evolve. One of the more mature, actively maintained multimodal vector stores, still shipping monthly.
CLI
- Entire CLI
- Entire CLI hooks into git and records what an AI coding agent actually did during a session — the prompts, the files touched, the tool calls — then indexes that transcript alongside the commit it produced. You can resume a session from any earlier checkpoint, or hand a teammate the full "why did the code change" story instead of just a diff. It shipped v0.10.2 on August 19, six days before this list went out, and has already logged nearly 8,000 commits of its own. Turns agent sessions into a searchable part of git history instead of a conversation that evaporates.
AI Agents
- ai-memory
- ai-memory gives coding agents a memory that survives past the session and past the tool: quit Claude Code mid-task, open Codex in the same directory, and it picks up the architecture decisions, failed approaches, and open questions without you re-explaining any of it. It ships native binaries for macOS, Linux, and Windows via WSL2, and it's barely three months old — created in May, already past 4,600 stars. Solves the specific, annoying problem of losing context every time you switch coding agents mid-task.