I write about papers and books I'm reading, and things I'm building, on my blog here.
Most recently, I've been working on AI math. I solved the 33 year old Erdos problem #81 and a number of conjectures alongside longtime collaborator N0zoM1z0
My personal research has been in agents, RL environments and preference learning if you would like to collaborate on research. You can reach me on X or my email here..
- REA: Reverse Engineer Anything
- Reverse engineer anything with agents, from app behavior down to native binaries.
- Awesome Reverse Engineering
- Curated reverse-engineering tools and learning resources for binaries, apps, firmware, file formats, and protocols.
- Jacobian
- Pure mathematics for agents: search for examples and counterexamples, compute exactly, and independently check what a result proves.
- Preference - Research infrastructure for trading agents.
- LeanToken
- Code intelligence for agents: find the code that matters and keep your context window and tokens lean.
- flameox
- Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
- GitContribute
- Contribution research for agents: check repository guidance, related work, code context, and validation before writing a patch.
- SmokingGun
- Optimization evidence for agents: find complexity hotspots and test whether a proposed change is worth making.
- gstack
- I wrote the first PR (not from Garry)
- Reef
- Continual learning infrastructure for agents: preserve receipts across overlapping runs, cache commit history for constant-time training bookkeeping, detect record conflicts after compaction, and make training and rollback settlement safe to retry.
- AgentEnv Framework (Scale AI)
- Task and storage performance: avoid quadratic response-history diffs, reduce implicit scheduler graphs, reuse journal keys, batch Explorer reads, and index latest-version lookups.
- Triton
- Autotuner fractional top_k pruning fix, unsigned tl.sum dtype promotion doc correction, invalid k value rejection in tl.topk, and tuple-target support in list comprehensions.
- NeMo labs-molt
- Vectorized default REINFORCE returns for faster RL training.
- slime
- Vectorized REINFORCE++ discounted returns for faster LLM post-training RL.
- Model Context Protocol
- TypeScript SDK packaging fix for CommonJS validator exports and Ajv.
- RepoPrompt
- Context Builder model selection, agent-mode and tool-card reliability, workspace model drift cleanup, test harness hardening against hangs and waiter leaks, hosted-app/XCTest sharding, superseded PR CI run cancellation, pending context preservation through routing, focused-build cost reporting, diff chunk reconstruction without repeated array shifts, replace-all matching with reused full-file indexes, and bounded LRU diff caching.
- DeepSeek Agents
- DeepSeek integration guides.
- Flash Linear Attention
- Split attention decode output offset fix, concurrent LSE store prevention in value-split attention, value-split Wall backward with local deltas, Wall autotuning reuse across length buckets, centralized generation errors for unsupported cache strategies, DPLR safe-gate full token stride offsets, tuple-form legacy cache round-trip restoration, explicit cu_seqlens preservation with attention masks, GQA head validation, strided decoding input handling, and context-parallel token partition validation.
- TileLang
- HIP atomic load/store and return_prev vector atomic add support, CUDA atomic add memory-order mapping to acquire PTX, FP8 E4M3 special-value decoding, BufferRegion destinations in atomic_addx2 return_prev, T.transpose axis-swap correctness, autotuner cache reuse prevention across different outputs and values, shared-TMEM buffer pointer type checks before dereference, unsupported TMA atomic add dtype rejection, lane-wise HIP vector Select lowering, flat CUDA include discovery for NVRTC, and symlink-aware nvcc discovery.
- Qwen Code
- Config validation for fractional session and tool-call limits, max-token continuation rollback, MCP read-only auto-approval trust enforcement, descendant termination after discovery timeout, restrictive permission path canonicalization, and malformed tool result display resilience.
- Executor
- Deno subprocess stdin failure handling, stalled GraphQL invocation timeout, failed policy and OAuth app mutation surfacing, and provider migration row ID collision prevention.
- Hermes
- Cron scheduler double-execution prevention and credential-pool failover fix for provider auto-detection.
- x-JEPA
- Correctness fixes across the JEPA training and planning stack: Wasserstein midpoint targets, unimodal Beta policies, goal/state token alignment, SIGReg dtype handling, and discrete planning action indices.
- OpenClaw
- Shared phone identity canonicalization, Signal/iMessage routing fixes, WhatsApp normalization coverage, stale-chunk recovery unblock, subagent registry resurrection prevention, heartbeat responsiveness with large commitment queues, and qa-lab transport teardown before gateway shutdown.
- Mooncake
- RDMA completion resource validation, cross-thread SHM relocation mapping lifecycle, and NVMe-oF task batch rejection with cuFile completion correlation.
- AgentENV
- OverlayBD fix to reject truncated mapped reads.
- gbrain
- Trajectory regression fix: stop negative metrics from inverting the signal.
- Orca
- Onboarding fix to skip Computer Use setup when the macOS helper app is unavailable.
- Harbor
- Make create-task and publish skills load in Codex.
- Vibe-Trading
- Codex OAuth init/default-model fixes and CLI resume prompt input preservation.
- SymPy
- Fix Pell-Gordon subresultant crashes when polynomials share a common factor by terminating on a zero remainder.
- Cordis
- Correct symbol and prototype-named event dispatch, and settle concurrent async interval reads in order.
- Magpie
- Fallible JSON parsing across FFI, runtime, codegen, semantic analysis, and web lowering.
- gitcrawl
- Portable store worktree verification to prevent local database misclassification under stray Git metadata.
- Crabbox
- RunPod fix to make new pods accept the configured SSH key.
- UniRL
- Sampling fix to make omitted autoregressive top-k unrestricted, unified loss reporting alignment with backward scaling, finite normalization for invalid rewards, per-prompt VideoAlign video paths, and bounded ZeroMQ operations during bucketed weight transfer.
- HarnessGym
- Benchmark correctness fixes for result parsing, activation reporting, and artifact quarantine.
- Qwen Code Docs
- Failed batch file reporting and output validation with bounded prose chunks in the docs translator, plus navbar locale preservation behind deployment base paths.




