Can symetric encryption alone ensure integrity without a signature?

If I have something that is symmetrically encrypted, like my KeePass or Notesnook database, could someone who DONT has the encryption password add something to the database, without it being obvious?

For kdbx — no if there is no bug in client: KDBX File Format Specification - KeePass

Generally for symmetric encryption — yes, until AEAD is presented.

There’s a few layers here. If you’re using a naive stream cipher like ChaCha20 or symmetric cipher in counter-mode (e.g. AES-CTR) , then known plaintext attacks will allow changing the known part to whatever the attacker wants with

PT  ⊕ CT = KS
PT' ⊕ KS = CT'

where PT is the plaintext, CT is the ciphertext, KS is the keystream, PT’ is the attacker chosen plaintext and CT’ is the attacker generated replacement ciphertext.

This is not possible with e.g. AES-CBC where next ciphertext block depends on previous one.

But modern symmetric ciphers are practically never naive, but either paired with a message authentication code (MAC) algorithm (e.g. HMAC-SHA256), or the symmetric cipher is a so-called AEAD construction where the MAC algorithm is part of the encryption (e.g. AES-GCM or ChaCha20-Poly1305).

MACs are much more common integrity verification method for personal file verification than digital signatures like RSA/ed25519, but as MAC algorithms also “sign” and “verify” the ciphertext (called tag in the context of MACs), we risk getting lost in the weeds on terminology.

As per KDBX 4 - KeePass the KDBX4 format that e.g. KeepassXC and Keepass2 use, uses HMAC-SHA256 for header and data authentication (under encrypt-then-MAC config which is best practice).

How does Notesnook encrypt my notes? | Notesnook Help says they’re using libsodium’s XChaCha20-Poly1305-IETF which is just about the best choice for symmetric crypto out there.

So for those two programs, it’s not possible to add more data, because the attacker would have to have the key, and for modification of existing data, the MACs prevent it.

Very well formed post, followed along nicely. To elaborate further, the stream cipher argument is one of two core goals of cryptography: indistinguishability (IND) and nonmalleability (NM). The argument present is against nonmalleability, and so the goal + threat model would be NM-KPA (known plaintext attacks). Attackers have ciphertext and associated plaintexts. Naive ciphers, as said, fail this attack. And with AEAD, you also prevent other classes of attacks like padding oracles and whatnot, as CBC mode can be vulnerable to that if not attached with a MAC (or say GCM is used or some other AEAD construction).

And correct me if I’m wrong, still sinking teeth into cryptography.

No I think you got it. I was thinking about the padding oracles wrt CBC but I’m unsure if user is actually decrypting the data over and over again while attacker makes changes to it, so it probably doesn’t count as a valid decryption oracle. But this thread has a sniff of it looking for generalized advice, so it should indeed be stated that CBC isn’t valid integrity mechanism; you always need a MAC.

So KeePass ensures integrity?

Does known plaintext means that the attacker knows the unencrypted version of something in the database?

How can he replaces it with something and encrypts it to my password, without having the password?

Yes.

KPA threats is that the attacker has the associated plaintext alongside the ciphertext. It’s threat modeling for cryptography - we assume that the attacker has access to both for cracking. You can choose different threats.

And usually such attacks are based an exploiting the malleability of a cipher (at which lets assume this is secure) or exploiting other aspects of the system and recover the private key (at which the attacker has full access). Don’t fret too hard, KeePass has done a lot of effort to protect against these attacks.

It doesn’t mean that they know secret values obviously. But e.g. during WW2, allies guessed some ciphertexts started with the word “Wetter”, because of the roughly even length and punctuality. (Note: Enigma wasn’t a stream cipher). Often you can see part of the plaintext via the source code if the plaintext contains e.g. JSON key-value pairs:

{
    first_name: "John",
    last_name: "Smith",
    age: 69
}

Even if you don’t know the name, you know it starts with {first_name because that’s part of the application code. It’s there for all users. When {first_name: is XORed with hex value 00161e1c1610000000000000, that part changes to {pwned_name:, which in turn causes error when the program can’t find the key first_name in the data structure.

So NM-KPA-security refers to preventing these kinds of situations, where attacker has some prior knowledge about what it would say. They obtain that information some other way, by reading the source code, and sometimes (like with Enigma) via educational guessing.

But the attacker doesn’t have the secret parts, that is, your notes or passwords so stream cipher’s unpredictability property means knowing the fixed parts doesn’t help them at all. And like was said above, the MACs prevent changing the ciphertext, and forging a tag is as hard as breaking the encryption itself.

That’s the neat part about modern cryptography, the attacker won’t. The weakest link to break the cryptographic guarantees is to break your password. So there’s no

Ciphertext -> Supercomputer -> Plaintext or
Ciphertext -> Supercomputer -> Key

There’s only

Password1+salt -> Password hashing function -> Key candidate1
Password2+salt -> Password hashing function -> Key candidate2
...

And with modern memory-hard hash functions like Argon2, that part is super slow, even with a super-computer provided the password itself is strong enough:

To maximize security, you’ll want to use long, random, unique passwords just about everywhere, and then protect the Keepass database with single strong master password. Today, a sufficiently strong master password is around 100 bits. KeepassXC has entropy calculator that’s fairly robust. Just make sure you don’t forget it, and have at least one hard-copy of the password hidden in some really safe place. Also, naturally, make sure to have enough backups of the password database itself.

I won’t explain it better than others here. But just in case: is this a theoretical question about encryption (which is cool), or is there something specific bothering you?

Beyond pure cryptography there are other things affecting the experience. For example, KeePass and some apps in that family implement various merging strategies. So if the same database is open in different apps, even on different machines, adding entry to the synced database on one machine can be synced and merged by another. Obviously the user on the first machine doesn’t need to know the secret if they have access to the opened Keepass with decrypted database.

Do they include signatures or how do they do it?

Are most encryption applications resistant against such attacks?

In my case these application are mostly:

  • KeePass
  • Notesnook
  • Cryptomator
  • Veracrypt and LUKS encryption tools

So even if the content is not signed and just encrypted with AES and the attacker knows some unencrypted version of parts of it, he cannot change the parts that he knows?

“AES” alone is not enough information to determine that. Again, with unauthenticated AES-CTR it’s entirely possible, with unauthenticated AES-CBC, most likely not. When I said modern cryptography, I assumed it was correctly implemented, i.e., deployed with the authentication mechanism. AES-CTR+HMAC-SHA256 is what the implementation should be using and that’s safe.

I’m getting the impression you’re trying to learn basics as loose facts as opposed to as a system that makes sense. Here’s a fantastic 101 book to get started with all this: https://raw.githubusercontent.com/crypto101/crypto101.github.io/master/Crypto101.pdf

I would actually recommend anyone here to read it, maybe even have a separate thread for it on PG forums to help each other with something that isn’t intuitively explained.

Short answer: yes.

Long answer: threat modeling in cryptography is a different beast than what we do here. The mentioned threat models are largely dependent on how they are utilized within the protocols themselves. And we are talking about the primitives (building blocks) of a larger protocol.

That said, any modern applications is built with a lot of the lessons of the past. Tools designed for cryptography (password managers, file encryptors) put a ton of energy in ensuring the protocols and primitives they used are done correctly.

The attacks mentioned like padding oracle attacks should be well known by any cryptographer worth their merit, and they will build the systems to prevent such basic footguns. But of course there is always risk something wasn’t implemented correctly.