I tried, but I don’t understand the LLM related jargon.
Most of the answers here are pretty useless because people don’t really know how LLMs work and answers are mostly determined by their sometimes obtuse beliefs. LLM generated output is not random in the same sense of randomly generated passwords which are a sampling of a predefined set of elements where each element has the same probability of being chosen, but it is also not deterministic and definetly not determined by “rules” like someone is saying. LLM generated output IS random but it is a sampling from a distribution of elements where each element have vastly different probabilities. The probabilities each element gets is determined by different factors: the internal weights of the LLM, the context of the element (token) that is being chosen at each moment during the text generation, and the decoding parameters/method (e.g. top-k, top-p, beam search, greedy). The weights of an LLM are not something directly interpretable by humans, but they may be the factor that most undermines the objective of the password generatiom task because it makes things quite predictable, BUT the context can be havily tweaked (through the prompt) and the decoding parameters can also be tweaked. People think LLMs generate a certain type of text but that’s because production models are all using pretty similar decoding parameters, which are the ones that produce meaningful text. However there are certain parameters (mostly temperature, I think) that can make the decoding process generate nonsense random text because the decoding parameters can be set to extreme values which in practice result in every element having the same probability. Therefore, there could be middle points where the generated text has much more randomness than the usual LLM output but still preserving some meaning in order to be (maybe) easier to memorize than diceware. So, in the end, it is up to the research to determine if tweaking both decoding and prompts in very different ways (which is what the paper linked above is talking about) can lead to outputs that have sufficient entropy so that they can be used as strong passwords.
Having said all that, I think teaching people to use diceware and get used to it is better and doing research in the LLM method is not worth it, but who knows.