I build AI tools and voice interfaces. I’m a software engineer at Kardium; previously Amazon.
Latest model comparison
More reasoning can cost more without improving the coding score. Faded settings are reference points.
Task cost includes tokens consumed. Per-token prices alone miss that.
DeepSWE v1.1 · official launch results · checked October 1, 2026.
Sol and Astra · Sep 29 · Luna · Sep 22
Routine and complex are starting choices: cheapest within 10 and 3 score points of the best, respectively. Test on your tasks. No comparable ultra results are published.
I built Live Transcriber for real-time speech transcription in Python. Source · 15K+ downloads.
Health Communication Reviewer checks claims against evidence and flags wording audiences may misinterpret.
I fixed a circular dependency in voice streaming in the OpenAI Agents SDK. Merged.