2024/
14 pages · Updated May 1, 2026
Pages
- 2024/06/02/main-swe-bench-html.html
- LLMs are bad at returning code in JSON
- 2024/05/22/swe-bench-lite-html.html
- Linting code for LLMs with tree-sitter
- Coding with Llama 3.1, new DeepSeek Coder & Mistral Large
- Claude 3 beats GPT-4 on Aider’s code editing benchmark
- Sonnet is the opposite of lazy
- Aider has written 7% of its own code (outdated, now 70%)
- Drawing graphs with aider, GPT-4o and matplotlib
- GPT-4 Turbo with Vision is a step backwards for coding
- Sonnet seems as good as ever
- The January GPT-4 Turbo is lazier than the last version
- Aider in your browser
- A draft post