On this page:
4.1 How I learned to stop worrying and love the LLM
4.2 Independent learning
9.1

4 Language models🔗

Probabilities come from all sorts of places. An interesting source of probabilities these days are large language models! What happens if we put them in our probabilistic programming languages?

Recently, a new kind of large language model came out that people are calling a "System-1 models", which focuses not on the free-form generation of text but rather the output of answers to questions with calibrated uncertainties. The first instance of this model was Jev, but since then several open source models have been released including kev and SemIf (formerly amusingly named "OpenJev").

These System-1 models work in a way that is very compatible with the probabilistic programs we’ve been exploring so far, which makes for some interesting capabilities.

Concretely, they support the following user interface for making probabilistic choices:

  • You provide a state, which is a string describing the state of the world. For example, you might say "it is rainy outside".

  • Next, you provide a question as a string. For example, "should I bring an umbrella?"

  • Next, you provide a list of possible answers as strings, for example "yes" or "no".

  • Then, the System-1 model will take this information in and return a probability of each of the answers being true.

Let’s see an example of programming with this. Here is a link to a small Roulette file that enables interaction with Jev; you’ll need to provide your own interaction key (we are hoping to make it able to call open models soon too, but haven’t gotten that going yet). To make it simple, we’ve provided a function called q-and-a-prob that takes 3 arguments: a state, a question, and a list of possible answers, and then returns a categorical distribution over the result:

> (define weather (q-and-a-prob "it rained today."
                                "will it rain tomorrow?"
                                (list "yes" "no")))
> weather

┌─────┬───────────────────┐

│Value│Probability        │

├─────┼───────────────────┤

│yes  │0.16000000000000003│

│no   │0.84               │

└─────┴───────────────────┘

Clearly, the above code is asking of an an absurd question, and some probability will certainly come out. But what do these probabilities mean? Broadly speaking, there are two ways of understanding probabilities as we have been using them so far:

  • Relative frequencies of outcomes. This is the frequentist interpretation of probabilities. A good example of this is coin flipping: we say a coin has probability 1/2 of landing on heads if, after flipping the coin many times, the ratio of heads to tails outcomes trends towards 1/2. This interpretation of probability has the advantage of a clear objective interpretation, but it suffers because it cannot handle events that happen only a single time, or events for which there is no way to run many trials.

  • Degrees of belief, or Bayesian probability. A Bayesian probability can be used as a measure or quantify one’s own confidence in an outcome. In this interpretation, one can simply assign a number to any event, and more likely events should get higher numbers. There is a rich literature and intellectual tradition on the rational ways of choosing Bayesian probabilities: Wikipedia has a good summary.

So, the interpretation of the above probability falls squarely in Bayesian camp: it is a probability coming from an LLM, and it’s up to the user to decide how to interpret that number. But, we can nonetheless program with it, and hopefully we can use it to do useful things despite its seeming arbitrariness. For example, suppose we don’t trust the LLM: we think it is right only 70% of the time, and the other 30% of the time it acts randomly. Then, of course we can write a probabilistic program that expresses this skepticism:

> (if (flip 7/10)
      weather
      (if (flip 1/2) "yes" "no"))

┌─────┬──────────────────┐

│Value│Probability       │

├─────┼──────────────────┤

│yes  │0.262             │

│no   │0.7380000000000001│

└─────┴──────────────────┘

This is fine, and lets us temper how much we trust whatever probability the LLM happens to give us. But, what would be better is if we could refine the probability that the LLM gives us by integrating it with other information.

With apologies to Dr. Strangelove

4.1 How I learned to stop worrying and love the LLM🔗

When should we trust what the LLM says? Clearly its probabilities are sometimes bogus and it’s sometimes wrong. But, other times it might be right, and it’s useful to know when that is happening too. Let’s try to set up a Roulette program that answers the question of how much should I believe the output of the LLM. One way to to combine the output of the LLM with other kinds of inputs. Let’s go back to the medical diagnosis example from earlier, and see how we can extend it with an LLM.

This time, suppose a patient comes into hospital. The doctor thinks they might have the flu, so they give them a flu test with a false positive rate of 20% and a false negative rate of 1%. On average, 1/20 patients have the flu. This time, the doctor also asks the LLM, given the patient’s verbal history, what the chances are that the patient has a flu.

The doctor has some trust in the LLM, but not a lot of trust. If the LLM is very confident (i.e, outputs a probability of 90% or higher either for having the flu or not having the flu), then the LLM tends to have the right diagnosis: the test in this case has a false-positive rate of 1% and false-negative rate of 0.1%. But, if the LLM is less confident, then it is much less reliable: it has a false-positive rate of 30% and false-negative rate of 15%.

Let’s model this situation to compute the probability that the patient has a flu:

We will see in later lectures how the parameters can be learned from data.

> (define (run-test has-flu? false-positive-rate false-negative-rate)
      (if has-flu
          (flip (- 1 false-positive-rate))
          (flip false-negative-rate)))
> (define has-flu (flip 1/20))
> (define flu-test-result (run-test has-flu 1/100 1/5))
> (define llm-query
      (q-and-a-prob (string-append "The patient presented with a mild fever, mild cough, and complaining"
                                   "of aches and pains. They are not taking any medications, and are"
                                   "otherwise healthy.")
      "Do they have the flu?"
      (list "yes" "no")))
> (define llm-result
      (if (> ((infer llm-query) (set "yes")) 0.9)
          (run-test has-flu 1/100 1/1000)
          (run-test has-flu 3/10 3/20)))
> (observe! flu-test-result)
> (observe! llm-result)
> has-flu

┌─────┬───────────┐

│Value│Probability│

├─────┼───────────┤

│#t   │231/421    │

│#f   │190/421    │

└─────┴───────────┘

We see above the use of a new Roulette feature: the infer keyword. This lets us, at runtime, compute the probability of some event happening in the program and manipulate those probabilities. We see in the above example that probabilistic programming languages can provide a good glue layer for rationally combining real-world information – such as the result of tests – with uncertain information coming from language models.

4.2 Independent learning🔗

  • Suppose you don’t trust the output of the language model very much and want it to have less impact on the final diagnosis for the patient. Update the above model to place less weight on the LLM result. How do you do this? Is there more than 1 way?