# Does this make sense? A passphrase generator based on a language model

**URL:** https://discuss.privacyguides.net/t/does-this-make-sense-a-passphrase-generator-based-on-a-language-model/41119
**Category:** Questions
**Tags:** please-eli5
**Created:** 2026-10-01T16:37:52Z
**Posts:** 22

## Post 1 by @NotVeryPrivateOnThisAccount — 2026-10-01T16:37:52Z

I was just thinking: LLMs can generate coherent sentences, and coherent sentences are easier to remember.

Would it be possible to ‘seed’ some kind of language model with randomness from a CSPRNG or true randomness source, and generate a pass-sentence, or pass-paragraph, that is above or equal to a target entropy?

It might be easier if a language designed for computers as well as humans was used, such as [Lojban](https://www.lojban.org/) or [Toaq](https://toaq.net/).

---

## Post 2 by @anonymous644 — 2026-10-01T16:41:31Z

An easier solution is to just use a resource like [strongphrase.net](http://strongphrase.net) which already has a passphrase generator

---

## Post 3 by @Valynor — 2026-10-01T16:47:29Z

I would say no: whatever comes out of a LLM is not _random_.  
You could argue it’s most likely to be statistical average.

---

## Post 4 by @fria — 2026-10-01T16:49:09Z

It’s the opposite of what you want. They generate the most likely string of words based on the input, so you would be much more likely to get the same output given the same input.

---

## Post 5 by @satchlj — 2026-10-01T17:22:08Z

> **[LLM-Generated Passphrases That Are Secure and Easy to Remember](https://aclanthology.org/2025.findings-naacl.290/)**
>
> Jie S. Li, Jonas Geiping, Micah Goldblum, Aniruddha Saha, Tom Goldstein. Findings of the Association for Computational Linguistics: NAACL 2025. 2025.

---

## Post 6 by @ph00lt0 — 2026-10-01T17:23:21Z

Why would you do this?

We have good ways of making random passwords s, this seems very counter intuitive

---

## Post 7 by @NotVeryPrivateOnThisAccount — 2026-10-01T18:54:10Z

Looks interesting. Thanks for sharing. Good to know I’m on the right track.

I’m tempted to mark it as the solution already, but I should probably read it first.

---

## Post 8 by @NotVeryPrivateOnThisAccount — 2026-10-01T18:57:22Z

I said in the OP: it would result in passphrases that are easier to remember.

Maybe it’s not _that_ useful, but it’s still a fun thought.

---

## Post 9 by @PaleCrow55 — 2026-10-01T19:58:49Z

> [@NotVeryPrivateOnThisAccount](#):
>
> passphrases that are easier to remember

If that’s what you want, then just use StrongPhrase, as previously mentioned.

---

## Post 10 by @anonymous644 — 2026-10-01T20:25:22Z

Interesting. So it does not look like the researchers even attempt to generate true randomness but instead raise entropy “through in-context examples and generation through a new top-q truncation method”

---

## Post 11 by @ph00lt0 — 2026-10-01T20:42:11Z

Yeah but like how hard is it to remember six random words.

---

## Post 12 by @blibly — 2026-10-02T10:35:28Z

In my experience, the hard part is typing the damn thing.

---

## Post 13 by @fria — 2026-10-02T10:59:47Z

Yeah it’s useful to have biometrics as a quicker authentication method after you type the password in once. GOS lets you set a separate, shorter PIN to use with biometrics which is really nice.

---

## Post 14 by @blibly — 2026-10-02T11:31:44Z

Doesn’t help if you do your important stuff on a computer like me. I had to shorten my computer’s password from five words to two.

---

## Post 15 by @dataSamurai — 2026-10-02T12:17:24Z

LLMs have statistically low entropy. Today security standards are way more high than just “it looks random so we’ll take it”. Some junior programmers still believe to this day that time (Unix timestamp) is a good randomizer, e.g. as a salt. It is not. You can predict or guess that.

Same with LLM, even today, when they are still a pretty new thing, there are papers on how you can predict what an LLM is going to give you based on the prompt.

@satchlj The issue with this paper is that LLMs are not deterministic. Even if the authors believe it’s going to produce high-quality, secure passwords, will it happen in every case? No.

> [@ph00lt0](#):
>
> We have good ways of making random passwords

Yeah, at this point I would rather spend time on creating secure passwords before quantum computers hit the market.

> [@NotVeryPrivateOnThisAccount](#):
>
> it would result in passphrases that are easier to remember.

If that’s the only upside, why even bother?

One more thing coming up to my mind is that a model like that would have to be local. We can’t allow for LLMs to produce our passwords, that’s counter-intuitive given that LLMs log and digest everything they do.

---

## Post 16 by @NotVeryPrivateOnThisAccount — 2026-10-02T13:28:13Z

> [@dataSamurai](#):
>
> Today security standards are way more high than just “it looks random so we’ll take it”.

I am aware of that. I was talking about seeding a model with the right kind of randomness, because they are random already. I definitely wasn’t talking about just asking ChatGPT for a password lol.

> [@dataSamurai](#):
>
> One more thing coming up to my mind is that a model like that would have to be local.

Of course.

---

## Post 17 by @Valynor — 2026-10-02T16:06:44Z

> [@NotVeryPrivateOnThisAccount](#):
>
> was talking about seeding a model with the right kind of randomness, because they are random already.

LLMs are in a way the opposite of random. If you take something already random and apply _rules_ to it it’s not random anymore.

---

## Post 18 by @Colter — 2026-10-02T16:25:42Z

1. Literally every model tells you that you should not put any secrets into it, and generating secrets with it, falls into the same category
2. Even if you have an local AI model: AI is not random but produces the most likely result
3. Even if you solve 1 and 2, it’s still a bad idea because the number of possible sentences is by magnitudes smaller then the number of word combinations in general.

---

## Post 19 by @NotVeryPrivateOnThisAccount — 2026-10-03T09:20:11Z

> [@Colter](#):
>
> 1. Even if you solve 1 and 2, it’s still a bad idea because the number of possible sentences is by magnitudes smaller then the number of word combinations in general.

> [@NotVeryPrivateOnThisAccount](#):
>
> and generate a pass-sentence, **or pass-paragraph,**

Then again, that might become difficult to type and/or remember again. Imagine having to reproduce an entire paragraph phrased in a specific way each time. Almost as bad as remembering a password with random letters and numbers and symbols. Then again, I guess it depends on the person.

---

## Post 20 by @NotVeryPrivateOnThisAccount — 2026-10-03T09:21:11Z

How many people responding to this idea have fully read the paper @satchlj posted?

---

## Post 21 by @NotVeryPrivateOnThisAccount — 2026-10-03T09:37:59Z

I tried, but I don’t understand the LLM related jargon.

---

## Post 22 by @CosmicOrca68 — 2026-10-03T10:10:17Z

Most of the answers here are pretty useless because people don’t really know how LLMs work and answers are mostly determined by their sometimes obtuse beliefs. LLM generated output is not random in the same sense of randomly generated passwords which are a sampling of a predefined set of elements where each element has the same probability of being chosen, but it is also not deterministic and definetly not determined by “rules” like someone is saying. LLM generated output IS random but it is a sampling from a distribution of elements where each element have vastly different probabilities. The probabilities each element gets is determined by different factors: the internal weights of the LLM, the context of the element (token) that is being chosen at each moment during the text generation, and the decoding parameters/method (e.g. top-k, top-p, beam search, greedy). The weights of an LLM are not something directly interpretable by humans, but they may be the factor that most undermines the objective of the password generatiom task because it makes things quite predictable, BUT the context can be havily tweaked (through the prompt) and the decoding parameters can also be tweaked. People think LLMs generate a certain type of text but that’s because production models are all using pretty similar decoding parameters, which are the ones that produce meaningful text. However there are certain parameters (mostly temperature, I think) that can make the decoding process generate nonsense random text because the decoding parameters can be set to extreme values which in practice result in every element having the same probability. Therefore, **there could be middle points where the generated text has much more randomness than the usual LLM output but still preserving some meaning in order to be (maybe) easier to memorize than diceware**. So, in the end, it is up to the research to determine if tweaking both decoding and prompts in very different ways (which is what the paper linked above is talking about) can lead to outputs that have sufficient entropy so that they can be used as strong passwords.

Having said all that, I think teaching people to use diceware and get used to it is better and doing research in the LLM method is not worth it, but who knows.
