One interesting consequence of the rise of LLMs: there's more demand for tools that handle untrusted input.
Arbitrary HTML+JS can be safely run in a browser. Lean can check an arbitrary proof.
These work really well with an LLM that can be wrong, but sometimes gives exactly what you want. Are there other tools in this family?
miniblog.
Related Posts
I've been experimenting with LLM autoresearch. I prompted Sol to make difftastic faster without changing output on the test suite, and log everything it tried:
https://github.com/Wilfred/difftastic/blob/c6e9c6fed4c71276a8b0e377c92695c3d929dc7f/PERF_RESEARCH_LOG.md
It found some interesting performance tweaks, although it's too tolerant of complexity.
Compiler error messages is such a deep topic. Even in rustc, a mature compiler that invests effort in diagnostic quality, I find interesting issues every few months.
(The majority have been fixed pretty quickly, so it's really rewarding to file issues.)
nREPL is a really interesting protocol for developer tools. It's extensible, but one of the basic operations is eval().
If your nREPL server doesn't support a given operation, you can just send an eval request to achieve the same result!