MD5, SHA-256 and bcrypt: Hashing vs Password Hashing
Drafted with AI assistance. Every command and code example was run and its output checked before publication. How guides are made
A general-purpose hash such as SHA-256 is built to be fast, which is exactly what you want for checksums, content addresses and signatures, and exactly what you do not want for passwords. A password hash such as Argon2id, scrypt or bcrypt is deliberately slow, uses a unique random salt for every password and has a cost setting you can raise over time, so that each guess costs an attacker real time and memory. Use SHA-256 (or SHA-3) to fingerprint data; use Argon2id to store passwords; use MD5 and SHA-1 only where nobody gains from a collision.
What does a cryptographic hash guarantee?
A hash function maps input of any length to a fixed-length digest. SHA-256 always produces 256 bits, written as 64 hex characters, and the same input always produces the same digest:
$ printf '%s' hello | sha256sum
2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824 -
$ printf '%s' Hello | sha256sum
185f8db32271fe25f561a6fc938b2e264306ec304eda518007d1764826381969 -
A cryptographic hash adds three security properties, the ones RFC 6151 lists as the design goals of message digests:
- Preimage resistance: given a digest, you cannot find any input that produces it.
- Second-preimage resistance: given one input, you cannot find a different input with the same digest.
- Collision resistance: you cannot find any two different inputs with the same digest. This is the hardest property to keep, and the first to fall.
What a hash does not guarantee is secrecy for guessable inputs. There is no key, so anyone can hash candidate inputs and compare. Hashing a short or predictable value, such as a password, an email address or a phone number, does not hide it.
Why are MD5 and SHA-1 considered broken?
Both have lost collision resistance in practice.
- MD5 (RFC 1321, 1992). The first MD5 collision pairs were published in 2004 and the techniques at EUROCRYPT 2005; RFC 6151 cites later work that finds a collision in 10 seconds or less on a 2.6 GHz Pentium 4. In 2008 researchers used MD5 collisions to create a rogue certification authority certificate trusted by all common browsers. RFC 6151 concludes that MD5 is no longer acceptable where collision resistance is required, such as digital signatures.
- SHA-1. In February 2017 Google and CWI Amsterdam announced the first practical SHA-1 collision (known as SHAttered; paper): two different PDF files with the same SHA-1 digest. Google put the cost at about 9.2 quintillion (263) SHA-1 computations. In 2020 the SHA-1 is a Shambles team computed the first chosen-prefix collision, where the attacker picks the start of both files, the kind of attack used against MD5 certificates. NIST announced in 2022 that SHA-1 should be phased out by 31 December 2030.
You can check the SHAttered result yourself. The two published PDFs are the same size, hash identically with SHA-1 and differently with SHA-256:
$ sha1sum shattered-1.pdf shattered-2.pdf
38762cf7f55934b34d179ae6a4c80cadccbb7f0a shattered-1.pdf
38762cf7f55934b34d179ae6a4c80cadccbb7f0a shattered-2.pdf
$ sha256sum shattered-1.pdf shattered-2.pdf
2bb787a73e37352f92383abe7e2902936d1059ad9f1ba6daaa9c1e58ee6970d0 shattered-1.pdf
d4488775d29bdef7993367d541064dbdda50d383f89f0aa13a6ff2e0894ba5ff shattered-2.pdf
Note what was broken: collisions, not preimages. No practical attack finds an input for a given MD5 or SHA-1 digest. That is why both are still acceptable for detecting accidental corruption, cache keys and de-duplicating trusted data, and unacceptable for signatures, certificates or anything where an attacker supplies the input.
SHA-2 and SHA-3: which should I use?
SHA-2 is a family defined in FIPS 180-4: SHA-224, SHA-256, SHA-384, SHA-512 and truncated variants. SHA-3 (FIPS 202) is a different design, based on the Keccak sponge construction, standardised in 2015 alongside SHA-2. NIST's SHA-1 announcement points users to either family. SHA-256 is the default choice: it is supported everywhere, including browsers' Web Crypto API. SHA-3 gives different digests for the same input, so the two are not interchangeable:
$ printf '%s' hello | openssl dgst -sha3-256
SHA3-256(stdin)= 3338be694f50c5f338814986cdf0686453a888b84f424d792af4b9202398f392
| Algorithm | Digest | Collision resistance today | Fine for |
|---|---|---|---|
| MD5 | 128 bits | Broken (collisions in seconds) | Non-adversarial checksums, cache keys |
| SHA-1 | 160 bits | Broken (practical collisions since 2017) | Legacy compatibility only |
| SHA-256 / SHA-512 | 256 / 512 bits | No known practical attack | Integrity, signatures, HMAC, content addressing |
| SHA3-256 | 256 bits | No known practical attack | Same uses as SHA-256 |
| Argon2id, scrypt, bcrypt | Varies | Not the point; designed to be slow | Password storage only |
The Hash Generator computes MD5, SHA-1, SHA-256, SHA-384 and SHA-512 of text at once, using the UTF-8 bytes of the input and showing lowercase hex. It does not offer SHA-3 or any password hash, because neither belongs in a quick text tool.
What are fast hashes good for?
- Checksums. A download page lists a SHA-256 digest; you hash the file and compare. This detects corruption and tampering, as long as the digest comes from a source you trust more than the file.
- Content addressing. The digest becomes the identifier: Git names objects by hash, and caches and build systems use digests as keys so that identical content is stored once.
- Message authentication with HMAC. To prove a message came from someone holding a shared secret, use HMAC (RFC 2104), not
sha256(secret + message), which is open to length-extension attacks with SHA-256. Compare signatures in constant time.
// Node.js
const crypto = require("crypto");
crypto.createHmac("sha256", "secret-key").update("GET /v1/orders?id=42").digest("hex");
// 'e1a97eea1e29560acfa470401213381473afd751ffe402d54d275084b3ca9f66'
# Python
import hmac, hashlib
hmac.new(b"secret-key", b"GET /v1/orders?id=42", hashlib.sha256).hexdigest()
# same value
The JWT Decoder verifies exactly this kind of HMAC-SHA-256 signature for HS256 tokens; How JSON Web Tokens work explains the format.
Why is SHA-256 wrong for storing passwords?
Because attackers do not attack the hash function; they guess the password. When a password database leaks, the attacker hashes candidate passwords from wordlists, leaked password lists and brute force, and compares. Every property in the first section still holds, and it does not matter: SHA-256 was designed to be computed quickly, on GPUs as well as CPUs, so each guess is cheap. The OWASP Password Storage Cheat Sheet says it plainly: fast hashing algorithms such as SHA-256 are not suitable for password storage because they allow attackers to perform large numbers of guesses quickly.
Human-chosen passwords also have little entropy. Eight random lowercase letters give 268, about 2.1 × 1011, combinations (37.6 bits); a password a person invented is far weaker than that because it is not random. A password hash cannot add entropy, but it can make every guess hundreds of milliseconds and many megabytes expensive instead of almost free.
What does a salt do?
A salt is a random value, unique to each password, that is hashed together with it and stored next to the result. It is not secret. Without salts, two users with the same password have the same hash, and one precomputed table (a rainbow table) cracks every account that uses a common password. With salts, the attacker must attack each hash separately. Salting a fast hash is not enough on its own, though: it stops precomputation, not fast guessing. Modern password-hashing libraries generate and store the salt for you.
bcrypt vs scrypt vs Argon2id: which should I use?
Use Argon2id where it is available. Argon2 won the 2015 Password Hashing Competition, and Argon2id, the variant OWASP recommends, is specified in RFC 9106, and is memory-hard: each guess needs a configurable amount of RAM, which limits how many guesses a GPU can run in parallel. The OWASP Password Storage Cheat Sheet gives these recommendations, quoted from its summary:
- "Use Argon2id with a minimum configuration of 19 MiB of memory, an iteration count of 2, and 1 degree of parallelism."
- "If Argon2id is not available, use scrypt with a minimum CPU/memory cost parameter of (2^17), a minimum block size of 8 (1024 bytes), and a parallelization parameter of 1."
- "For legacy systems using bcrypt, use a work factor of 10 or more and with a password limit of 72 bytes."
- "If FIPS-140 compliance is required, use PBKDF2 with a work factor of 600,000 or more and set with an internal hash function of HMAC-SHA-256."
| Algorithm | OWASP minimum (one option) | Memory-hard | Notes |
|---|---|---|---|
| Argon2id | m=19456 (19 MiB), t=2, p=1 | Yes | First choice; equivalent trade-offs listed, e.g. 46 MiB with t=1 |
| scrypt | N=2^17 (128 MiB), r=8, p=1 | Yes | RFC 7914; built into Node.js and Python |
| bcrypt | Work factor 10 | No | Legacy; input limited to 72 bytes |
| PBKDF2-HMAC-SHA256 | 600,000 iterations | No | RFC 8018; when FIPS-140 is required |
These are minimums. OWASP also advises benchmarking on your own servers and keeping a hash under one second, and raising the cost as hardware improves; store the parameters with each hash so you can rehash on the user's next login.
Examples: PHP and Node.js
PHP has Argon2id and bcrypt built into password_hash (Argon2 requires PHP built with Argon2 support). The output string contains the algorithm, parameters, salt and hash, so verifying needs nothing else:
<?php
$hash = password_hash($password, PASSWORD_ARGON2ID,
['memory_cost' => 19456, 'time_cost' => 2, 'threads' => 1]);
// $argon2id$v=19$m=19456,t=2,p=1$<salt>$<hash>
if (password_verify($password, $hash)) {
if (password_needs_rehash($hash, PASSWORD_ARGON2ID,
['memory_cost' => 19456, 'time_cost' => 2, 'threads' => 1])) {
// store password_hash($password, ...) again
}
}
PASSWORD_DEFAULT is still bcrypt; PHP 8.4 raised its default cost from 10 to 12. Node.js 20 has no built-in Argon2, but crypto.scryptSync is built in. Raise maxmem, or the OWASP parameters fail with a "memory limit exceeded" error:
const crypto = require("crypto");
const PARAMS = { N: 2 ** 17, r: 8, p: 1, maxmem: 256 * 1024 * 1024 };
function hashPassword(password) {
const salt = crypto.randomBytes(16);
const key = crypto.scryptSync(password, salt, 32, PARAMS);
return `scrypt$${PARAMS.N}$${PARAMS.r}$${PARAMS.p}$${salt.toString("base64")}$${key.toString("base64")}`;
}
function verifyPassword(password, stored) {
const [, N, r, p, salt, key] = stored.split("$");
const expected = Buffer.from(key, "base64");
const actual = crypto.scryptSync(password, Buffer.from(salt, "base64"), expected.length,
{ N: +N, r: +r, p: +p, maxmem: PARAMS.maxmem });
return crypto.timingSafeEqual(actual, expected);
}
In production, prefer a maintained password-hashing library over hand-rolled storage formats; the code above shows what such a library does.
The bcrypt 72-byte limit
bcrypt only uses the first 72 bytes of the password. In PHP 8.5, a bcrypt hash of 72 a characters followed by XYZ verifies against 72 a characters followed by anything else. With multi-byte UTF-8 characters, 72 bytes can be far fewer than 72 characters. OWASP's advice is to enforce a 72-byte maximum, or to pre-hash with an HMAC keyed by a pepper and Base64-encode the result before bcrypt; a plain unkeyed pre-hash lets attackers reuse leaked fast hashes ("password shucking").
What is a pepper?
A pepper is a secret shared by all passwords and kept outside the database, for example in a secrets vault or hardware security module. It is applied either before hashing or as an HMAC key over the finished password hash. If only the database leaks, say through SQL injection or a stolen backup, the hashes cannot be cracked without the pepper. OWASP describes it as defence in depth: it adds protection alongside salting and a slow hash, and does not replace either. The cost is operational: a pepper cannot be rotated without each user's password, so losing or changing it means password resets.
How much does the password itself matter?
A great deal. A slow hash multiplies the cost of each guess by a fixed factor; a random password multiplies the number of guesses needed exponentially. Sixteen characters drawn at random from 80 symbols give about 101 bits, which no hash setting needs to compensate for, while a common word is found early in any wordlist however slow the hash. The Password Generator creates passwords with the browser's secure random source and shows their estimated entropy. On the service side, NIST's SP 800-63B covers length requirements and checking new passwords against lists of compromised ones.
Checklist
- Integrity, signatures and content addressing: SHA-256 or stronger (SHA-512, SHA3-256).
- MD5 and SHA-1 only where no one benefits from a collision.
- Authenticating messages with a shared secret: HMAC-SHA-256, compared in constant time.
- Passwords: Argon2id (19 MiB, t=2, p=1 or stronger); else scrypt (N=2^17, r=8, p=1); bcrypt (cost 10 or more) only in legacy systems; PBKDF2-HMAC-SHA256 with 600,000 iterations where FIPS-140 is required.
- A unique random salt per password, stored with the hash (libraries do this).
- Parameters stored with each hash; rehash on login when they change.
- Optional pepper, stored outside the database.
- Never encrypt passwords for storage, and never use a fast hash, salted or not.
Frequently asked questions
Can a SHA-256 hash be reversed?
Not mathematically; no preimage attack on SHA-256 is known. But a hash of a guessable input can be found by hashing guesses, which is how online hash lookup sites work: they store the hashes of known words and leaked passwords.
Is hashing the same as encryption?
No. Encryption uses a key and is reversible by whoever holds it; hashing has no key and no way back. OWASP recommends hashing, not encryption, for passwords, because a service only needs to check a password, never to read it.
Is MD5 still OK for file checksums?
For detecting accidental corruption, yes. If an attacker could swap the file, no: MD5 collisions are cheap, so use the SHA-256 digest that most publishers now list.
Is bcrypt still safe to use?
OWASP still gives settings for it, with a work factor of 10 or more, but recommends it only for legacy systems where Argon2id and scrypt are not available. Mind the 72-byte input limit.
Why does the same password give a different hash every time?
Because password hashing functions add a new random salt each time. Verification reads the salt and parameters from the stored string and recomputes, so use the library's verify function instead of comparing two freshly made hashes.