Saturday, August 23, 2025

CL-JSON Ambiguity

When you are using a library, you will usually get best results if you avoid customizations, flags, and other non-default settings. Just use it out of the box with the vanilla settings. This isn't always true, but it is a good place to start. The default settings are likely to have been thoroughly tested, and you are more likely to encounter fellow travellers who are using the same settings. But if you have a good reason to change a setting, go ahead and do so.

I use the cl-json library for handling JSON. I encountered a bug in it. When using the default settings, certain objects won't round-trip through the decoder and encoder.

By default JSON arrays are decoded into lists and JSON objects are decoded into association lists. This is very convenient for us Lisp hackers, but there is an ambiguity inherent in this approach. If the value of a key in a JSON object is itself an array, then the decoded object, an association list, will have an entry that is a CONS of a key and a list. But a CONS of an element and a list is a larger list, so we cannot tell whether a list is an association list entry or simply a nested list. In other words, cl-json will decode these two different JSON objects into the same Lisp object:

{"a": [1, 2, 3]}    [["a", 1, 2, 3]]
(("a" . (1 2 3))) = (("a" 1 2 3))         

When encoding (("a" 1 2 3)), it is ambiguous as to whether to treat this as an association list and encode it as an object or to treat it as a nested list and encode it as a nested array.

Solving this problem is pretty easy, though. You customize the decoder to use vectors for JSON arrays and hash tables for JSON objects. Now there is no ambiguity because the encoder never encounters a list.

Unfortunately, it is a bit more clumsy to use. You cannot use list operations on vectors. But sequence operations will still work.


Saturday, August 16, 2025

Dinosaurs

What did the dinosaurs think in their twilight years as their numbers dwindled and small scurrying mammals began to challenge their dominance? Did they reminisce of the glory days when Tyrannosaurus Rex ruled the land and Pteranodon soared through the air? Probably not. They were, after all, just dumb animals.

Our company has decided to buy in to Cursor as an AI coding tool. Cursor is one of many AI coding tools that have recently been brought to market, and it is a fine tool. It is based on a fork of VSCode and has AI coding capabilities built in to it. One of the more useful ones (and one that is available in many other AI tools) is AI code completion. This anticipates what you are going to type and tries to complete it for you. It gets it right maybe 10-20% of the time if you are lucky, and not far wrong maybe 80% of the time. You can get into a flow where you reflexively keep or discard its suggestions or accept the near misses and then correct them. This turns out to be faster than typing everything yourself, once you get used to it. It isn't for everyone, but it works for me.

Our company has been using GitHub Copilot for several months now. There is an Emacs package that allows you to use the Copilot code completion in Emacs, and I have been using it for these past few months. In addition to code completion, it will complete sentences and paragraphs in text mode and html mode. I generally reject its suggestions because it doesn't phrase things the way I prefer, but I really like seeing the suggestions as I type. It offers an alternative train of thought that I can mull over. If the suggestions wildly diverge from what I am thinking, it is usually because I didn't lay the groundwork for my train of thought, so I can go back and rework my text to make it clearer. It seems to make my prose more focused.

But now comes Cursor, and it has one big problem. It is a closed proprietary tool with no API or SDK. It won't talk to Emacs. So do I abandon Emacs and jump on the Cursor bandwagon, or do I stick with Emacs and miss out on the latest AI coding tools? Is there really a question? I've been using Emacs since before my manager was born, and I am not about to give it up now. My company will continue with a few GitHub Copilot licenses for those that have a compelling reason to not switch to Cursor, and I think Emacs compatibility is pretty compelling.

But no one uses Emacs and Lisp anymore but us dinosaurs. They all have shiny new toys like Cursor and Golang. I live for the schadenfreude of watching the gen Z kids rediscover and attempt to solve the same problems that were solved fifty years ago. The same bugs, but the tools are now clumsier.


Tuesday, August 12, 2025

How Did I Possibly Break This?

It made no sense. My change to upgrade the Java Virtual Machine caused a number of our builds to stop working. But when I investigated, I found that the builds were failing in tsc, the TypeScript compiler. The TypeScript compiler isn't written Java. Java isn't involved in the tool chain. What was going on?

It turned out that someone pushed an update to a TypeScript library simultaneously (but purely coincidentally) with my Java upgrade. The code was written to use the latest library and our TypeScript compiler was past its prime. It barfed on the new library. Java was not involved in the least. It only looked causal because breakage happened right after I pushed the new image.


Monday, August 11, 2025

Why LLMs Suck at Lisp

In my experiments with vibe coding, I found that LLMs (Large Language Models) struggle with Lisp code. I think I know why.

Consider some library that exposes some resources to the programmer. It has an AllocFoo function that allocates a Foo object, and a FreeFoo function that frees it. The library his bindings in several languages, so maybe there is a Python binding, a C binding, etc. In these languages, you'll find that functions that call AllocFoo often call FreeFoo within the same function. There are a lot of libraries that do this, and it is a common pattern.

Documents, such as source code files, can be thougth of as “points” in a very high dimensional space. Source code files in a particular language will be somewhat near each other in a region of this space. But within the region of space that contains source code in some language, there will be sub-regions that exhibit particular patterns. There will be a sub-region that contains Alloc/Free pairs. This sub-region will be displaced from the center of the region for the language. But here's the important part: in each language, independent of the particulars of the language, the subregion that contains Alloc/Free pairs will be displaced in roughly the same direction. This is how the LLM can learn to recognize the pattern of usage across different languages.

When we encounter a new document, we know that if it is going to contain an Alloc/Free pair, it is going to be displaced in the same direction as other documents that contain such pairs. This allows us to pair up Alloc/Free calls in code we have never seen before in languages we have never seen before.

Now consider Lisp. In Lisp, we have a function that allocates a foo object, and a function that frees it. The LLM would have no problem pairing up alloc-foo and free-foo in Lisp. But Lisp programmers don't do that. They write a with-foo macro that contains an unwind-protect that frees the foo when the code is done. The LLM will observe the alloc/free pair in the source code of the macro — it looks like your typical alloc/free pair — but then you use the macro everywhere instead of the explicit calls to Alloc/Free. The LLM doesn't know this abstraction pattern. People don't write with-foo macros or their equivalents in other languages, so the LLM doesn't have a way to recognize the pattern.

The LLM is good at recognizing patterns, and source code typically contains a lot of patterns, and these patterns don't hugely vary across curly-brace languages. But when a Lisp programmer sees a pattern, he abstracts it and makes it go away with a macro or a higher-order function. People tend not to do that in other languages (largely because either the language cannot express it or it is insanely cumbersome). The LLM has a much harder time with Lisp because the programmers can easily hide the patterns from it.

I found in my experiments that the LLMs would generate Lisp code that would allocate or initialize a resource and then add deallocation and uninitialization code in every branch of the function. It did not seem to know about the with-… macros that would abstract this away.


Sunday, August 10, 2025

LLM in the Debugger

There is one more place I thought I could integrate the LLM with Lisp and that is in the debugger. The idea is to have the LLM have gander at the error and suggest a fix before you get dropped into the debugger as usual. The mechanism is pretty straightforward. You use the debugger hook to call the LLM with the error message and a backtrace and tell it you'd like a diagnosis and a suggested fix. You also tell it that it can use its tools to inspect the Lisp environment. Then you cross your fingers and hope that the LLM has some luck. At the very least you get a second opinion on your error.

I had some difficulty with this because SLIME smashes the *debugger-hook* on each interaction. Eventually I settled for a macro that you can wrap your code with. with-llm-debugger binds *debugger-hook* to the LLM debugger before running the body of the macro. Let's see it in action:

(defun my-average (list) (/ (reduce #'+ list) (length list)))

Now let's make the mistake of thinking it takes a &rest argument:

> (with-llm-debugger (my-average 2 3 4))
; in: GEMINI::WITH-LLM-DEBUGGER (MY-AVERAGE 2 3 4)
;     (MY-AVERAGE 2 3 4)
; 
; caught STYLE-WARNING:
;   The function MY-AVERAGE is called with three arguments, but wants exactly one.
; 
; compilation unit finished
;   caught 1 STYLE-WARNING condition
An unhandled error condition has been signalled: invalid number of arguments: 3

;;
;; **Troubleshooting a Lisp Argument Error**
;;
;; Okay, so I've got this Common Lisp error message: "invalid number of arguments:
;; 3".  And the backtrace pinpoints `MY-AVERAGE` as the culprit, called with the arguments
;; `2 3 4`. Hmm, that's pretty straightforward. The function `MY-AVERAGE` was called
;; with three arguments, but it wasn't expecting them.
;;
;; The next step, naturally, is to figure out what `MY-AVERAGE` *does* expect.  Since
;; I don't have direct access to the user's code, I'll need to query the Lisp environment
;; itself. The best way to do that is to use the `default_api.describe` function on
;; the symbol `MY-AVERAGE`. This should give me its definition, which will include
;; its expected argument list.
;;
;; Once I have that argument list, the fix should be obvious.  Either the arguments
;; being passed are wrong, or the `MY-AVERAGE` function needs a definition adjustment.
;;  I'm ready to proceed, query the environment, and get to the root of this.
;;
;; Processing function call: describe

The LLM has decided to call into Lisp. The describe function is a tool that is just a wrapper around the Lisp describe function. It takes a symbol and returns whatever the Lisp describe function returns.

;; **Understanding the `MY-AVERAGE` Function Call Error**
;;
;; Okay, so looking at this, it seems the problem is pretty straightforward. `describe`
;; clearly tells me that `MY-AVERAGE` is a compiled function and that it's expecting
;; a single argument, which it's calling `LIST`. I can see from the backtrace that
;; I attempted to call it with the individual numbers 2, 3, and 4. The error message
;; makes perfect sense now. I provided three separate arguments, but the function's
;; definition only allows for a single argument: a list.
;;
;; Essentially, I need to wrap those numbers in a list before passing them to the
;; function. Instead of calling it with `(MY-AVERAGE 2 3 4)`, I should have been using
;; `(MY-AVERAGE '(2 3 4))`. That way, the function receives a single list containing
;; the elements it expects. A classic case of mismatched expectations between function
;; definition and invocation! I'll update the calling code and that should resolve
;; the issue.
;;

So those are the LLM's “thoughts” on the error. Here is its diagnosis and conclusion:

The error "invalid number of arguments: 3" for `MY-AVERAGE`
occurred because the `MY-AVERAGE` function expects a single argument,
which should be a list, as indicated by its lambda list
`(LIST)`. However, it was called with three separate arguments: `2`,
`3`, and `4`.

**Diagnosis:**
The function `MY-AVERAGE` is defined as:

```commonlisp
(LAMBDA (LIST)
  (BLOCK MY-AVERAGE (/ (REDUCE #'+ LIST) (LENGTH LIST))))
```
;; Prompt Tokens:       4316
;; Thoughts Tokens:      227
;; Candidate Tokens:     223
;; Total Tokens:        4766
;; Prompt Tokens:       3959
;; Thoughts Tokens:      209
;; Candidate Tokens:      16
;; Total Tokens:        4184
LLM Debugger Response: 
This definition clearly shows that `MY-AVERAGE` is intended to take a single argument named `LIST`.

**Suggested Fix:**
To correctly use `MY-AVERAGE` to calculate the
average of 2, 3, and 4, you need to pass these numbers as a single
list argument.

Change the call from:
`(MY-AVERAGE 2 3 4)`

To:
`(MY-AVERAGE '(2 3 4))`

Not too bad, I guess. But I'd have to put it through its paces on real errors before I could say how useful it is.


Saturday, August 9, 2025

REPL + Prompt

I'm interested in LLMs and I'm into Lisp, so I wanted to explore ways to combine the two. Earlier, I wrote a pseudocode macro that uses an LLM to expand pseudocode into Common Lisp code. I also wrote an autodoc feature that uses an LLM to generate docstrings if you don't provide them yourself. These are two examples of Lisp calling into the LLM. Naturally, I wanted to see what would happen if we let the LLM call into Lisp.

We can provide Lisp “tools” to the LLM so that it can have an API to the client Lisp. Some of these tools are simply extensions to the LLM that happen to written in Lisp. For example, the random number generator. We can expose user interaction tools such as y-or-n-p to allow the LLM to ask simple y/n questions. But it is interesting to add Lisp introspection tools to the LLM so it can probe the Lisp runtime.

There is an impedance mismatch between Lisp and the LLM. In Lisp, data is typed. In the LLM, the tools are typed. In Lisp, a function can return whatever object it pleases and the object carries its own type. An LLM tool, however, must be declared to return a particular type of object and must return an object of that type. We cannot expose functions with polymorphic retun values to the LLM because we would have to declare the return type prior to calling the function. Furthermore, Lisp has a rich type system compared to that of the LLM. The types of many Lisp objects cannot easily be declared to the LLM.

We're going to live dangerously and attempt to give the LLM the ability to call eval. eval is the ultimate polymorphic function in that it can return any first-class object. There is no way to declare a return type for eval. There is also no way to declare the argument type of s-expression. Instead, we declare eval to operate on string representations. We provide a tool that takes a string and calls read-from-string on it, calls eval on the resulting s-expression, and calls print on the return value(s). The LLM can then call this tool to evaluate Lisp expressions. Since I'm not completely insane, I put in a belt and suspenders check to make sure that the LLM does not do something I might regret. First, the LLM is instructed to get positive confirmation from the user before evaluating anything that might have a permanent effect. Second, the tool makes a call to yes-or-no-p before the actual call to eval. You can omit this call to yes-or-no-p by setting *enable-eval* to :YOLO.

It was probably obvious to everyone else, but it took me a bit to figure out that maybe it would be interesting to have a modified REPL integrated with the LLM. If you enter a symbol or list at the REPL prompt, it would be sent to the evaluator as usual, but if you entered free-form text, it could be sent to the LLM. When the LLM has the tools capable of evaluating Lisp expressions, it can call back into Lisp, so you can type things like “what package am I in?” to the modified REPL, and it will generate a call to evaluate “(print *package*)”. Since it has access to your Lisp runtime, the LLM can be a Lisp programming assistant.

The modified REPL has a trick up its sleeve. When it gets a lisp expression to evaluate, it calls eval, but it pushes a fake call from the LLM to eval on to the LLM history. If a later prompt to is given to the LLM it can see this call in its history — it appears to the LLM as if it had made the call. This allows the LLM to refer to the user's interactions with Lisp. For example,

;; Normal evaluation
> (+ 2 3)
5

;; LLM command
> add seven to that
12

The LLM sees a call to eval("(+ 2 3)") resulting in "5" in its history, so it is able to determine that the pronoun “that” in the second call refers to that result.

Integrating the LLM with the REPL means you don't have to switch windows or lose context when you want to switch between the tools. This streamlines your workflow.


Thursday, August 7, 2025

Autodoc

The opposite of pseudocode is autodoc. If pseudocode is about generating code from a text description, then autodoc is about generating a text description from code. We shadow the usual Common Lisp defining symbols that take docstrings, such as defun, defgeneric, defvar, etc. and check if the docstring was supplied. If so, it is used as is, but if the docstring is missing, we ask the LLM to generate one for us. Your code becomes self-documenting in the truest sense of the word.

I have added autodoc as an adjunct to the pseudo system. It uses the same LLM client (at the moment Gemini) but the system instructions are tailored to generate docstrings. Here are some examples of source code and the generated docstrings:

(defclass 3d-point ()
  ((x :initarg :x :initform 0)
   (y :initarg :y :initform 0)
   (z :initarg :z :initform 0)))

;;; Generated docstring:
"A class representing a point in 3D space with x, y, and z coordinates."

(defconstant +the-ultimate-answer+ 42)

;;; Generated docstring:
"A constant representing the ultimate answer to life, the universe, and everything."

(defgeneric quux (a b)
  (:method ((a int) (b int))
    (declare (ignore a b))
    0)
  (:method ((a string) (b string))
    (declare (ignore a b))
    "Hello, world!"))

;;; Generated docstring:
"A generic function that takes two arguments of type int or string.

The int method returns 0, and the string method returns 'Hello, world!'."

(defmacro bar (a b)
  `(foo ,a ,b))

;;; Generated docstring:
"A macro that expands to a call to the function foo with arguments a and b."

(defparameter *screen-width* 640)

;;; Generated docstring:
"A global variable representing the width of the screen in pixels."  

(defstruct point
  (x 0)
  (y 0))

;;; Generated docstring:
"A structure representing a point in 2D space with x and y coordinates."

(defun foo (a b)
  (+ a b))

;;; Generated docstring:
"A function that takes two arguments a and b and returns their sum."

(defvar *current-foo* nil)

;;; Generated docstring:
"A global variable that holds the current value of foo, initialized to nil."

As you can see, the generated docstrings aren't bad. They describe the purpose of the class, constant, generic function, macro, global variable, and structure. The docstrings are not perfect, but they are better than nothing, which is what you start with.