📅 Day 229 — Hash Functions and Why They Are Not Encryption (Google Cybersecurity Certificate)
🔄 Topic
Studying hash functions in the Google Cybersecurity Certificate clarified something I’d been fuzzy on: hashing isn’t a weaker form of encryption, it’s a fundamentally different tool — one-way, irreversible, and built for proving integrity rather than hiding content.
🎯 Goal
Understand how hash functions differ from encryption, why hash collisions matter, and how rainbow tables and salting fit into real password-storage security.
🛠 What I Did
I worked through hashing’s origins, its core vulnerability, and the defenses built to address it.
Main areas covered:
- learned that a hash function produces a code — a hash value or digest — that cannot be decrypted, because unlike symmetric and asymmetric algorithms, hashing is a one-way process with no decryption key generated at all
- learned the core integrity use case: hash a file, store the hash, and any later change to that file — even a single line of code — produces a completely different hash value, making tampering detectable by comparison rather than by inspection
- connected this to non-repudiation: the concept that authenticity can’t be denied, and that hash functions are what make proven data integrity possible in the first place
- learned MD5’s history: created by Ronald Rivest at MIT in the early 1990s, producing a 128-bit value (a 32-character string) — and learned why it’s now considered weak
- learned about hash collisions: since hash functions map infinite possible inputs onto a finite set of fixed-size outputs, it’s mathematically guaranteed that some different inputs will eventually produce the same hash — MD5’s 32-character output space made this practically exploitable, since an attacker could potentially substitute one file for another that hashes identically, effectively forging authenticity
- learned the SHA family as MD5’s replacement: SHA-1 (160-bit, no longer considered collision-resistant), SHA-224, SHA-256, SHA-384, SHA-512 — all NIST-approved, and all except SHA-1 currently considered collision-resistant
- learned how password storage actually uses hashing: a server never stores a plaintext password, only its hash, and authentication works by hashing what the user just typed and comparing that to the stored hash — meaning even a full database breach doesn’t hand an attacker usable plaintext credentials directly
- learned about rainbow tables: precomputed dictionaries mapping common plaintexts to their hash values, letting an attacker with a stolen password database reverse-lookup weak or common passwords almost instantly
- learned salting as the defense: adding a random string of characters to data before hashing it, so that even identical passwords (“password” used by five different accounts) produce five completely different hash values, making a precomputed rainbow table useless against the salted database
🔗 Key Cybersecurity Connections
The distinction between “can’t be decrypted” (hashing) and “was decrypted with the wrong key” (a broken encryption attempt) matters practically: if an attacker steals a password database protected only by hashing, they don’t get a locked box they might eventually pick — they get output with no corresponding input recovery path at all, short of guessing. That’s a categorically stronger guarantee than encryption provides, which is exactly why passwords are hashed and not just encrypted.
Hash collisions are the sharp edge of an otherwise strong tool: MD5’s limited output space is precisely why organizations moving away from it isn’t paranoia, it’s math — a large enough space of possible collisions eventually becomes an exploitable identity-forgery risk. Salting is a small addition with an outsized effect specifically because it defeats precomputation: a rainbow table’s entire value comes from being calculated once and reused against many targets, and salting makes every target’s hash space unique, collapsing that reuse advantage to nothing.
🔍 Investigation Questions
- Is any password or sensitive data still hashed with MD5 or unsalted SHA-1 rather than a modern, salted, collision-resistant algorithm?
- Does a file-integrity check compare hash values, or does it rely on visual inspection or file size alone?
- Is salt generated uniquely per record, and long/random enough to resist precomputed-table attacks?
- Could a hash collision realistically be exploited against any integrity check currently in use, given the algorithm’s output size?
- If a password database were stolen today, would the hashing and salting in place actually slow an attacker down meaningfully?
🚨 Detection Opportunities
Checks for hashing and password-storage security:
- passwords or sensitive identifiers hashed with MD5 or unsalted SHA-1 instead of a modern algorithm
- identical plaintext values producing identical hash values across records (a sign salting is missing or reused)
- a file-integrity process relying on manual comparison instead of automated hash verification
- salt values that are short, predictable, or reused across records
- hash values exposed anywhere they could be harvested for offline rainbow-table or brute-force attacks
Example:
project=password-storage-review
signal=unsalted_hash_or_md5_in_use_for_credentials
risk_area=rainbow_table_vulnerability
triage=migrate_to_salted_modern_hash_algorithm
🧭 MITRE ATT&CK Techniques
Possible mapping for the risk this material addresses:
- T1552 — Unsecured Credentials (weak or unsalted hashing directly increases the risk of credential exposure if a database is compromised)
🗺 Visual Investigation Diagram
File or password input
↓
Hash function (one-way, no decryption key produced)
↓
Fixed-size digest stored (MD5 weak; SHA-2 family preferred)
↓
Later comparison: does the hash still match?
↓
Match = integrity intact / correct password
Mismatch = tampering detected / wrong password
↓
Defense against stolen hash database: salting defeats rainbow tables
⚠ Challenges
The part that took the most re-reading was hash collisions — understanding why they’re mathematically inevitable (finite output space, infinite possible inputs) rather than just accepting “MD5 is bad” as a fact to memorize. Working through the pigeonhole-style reasoning made the SHA family’s larger output sizes make sense as a real mitigation rather than an arbitrary improvement.
📚 What I Learned
I learned that hashing and encryption solve different problems and shouldn’t be mentally lumped together just because both involve transforming data — encryption protects confidentiality and is meant to be reversed by the right key holder, while hashing protects integrity and authenticity and is never meant to be reversed by anyone. Salting’s effectiveness against rainbow tables also taught me that a small, cheap addition (a random string) can defeat an entire category of precomputation attack by removing the reusability that made the attack economical in the first place.
➡ Next Steps
- Check whether any systems I maintain still use MD5 or unsalted hashing anywhere
- Practice generating and comparing SHA-256 hashes on real files to build the habit of verifying integrity by hash, not by eye
- Connect hashing back to the PKI material from yesterday — understand where digital signatures use hashing as a component
- Move on to the AAA framework and see how hashing underpins the authentication piece specifically
🧠 Reflection
The line that reframed this topic for me was realizing “can’t be decrypted” isn’t a limitation of hash functions — it’s the entire point. A tool that can’t be reversed is exactly what you want when the goal is proving something hasn’t changed, not hiding what it originally was.
🧩 Lessons Learned
What worked
Understanding hashing through what it’s structurally incapable of (reversal) rather than just what it’s used for, and connecting rainbow tables and salting as attack and defense that directly counter each other.
What broke
Nothing broke — this was concept study, grounded in a real historical case (MD5’s collision weakness) rather than a hands-on incident.
Why it mattered
Password and integrity protection depend entirely on choosing a hash algorithm and salting strategy that actually resist the attacks designed against weaker predecessors.
Fix / takeaway
Use modern, collision-resistant hash algorithms (SHA-256 or stronger), always salt sensitive hashed data uniquely per record, and verify integrity by comparing hash values rather than trusting appearances.
📈 Skill Progression Context
This supports my cybersecurity progression because hashing, collision resistance, and salting are foundational to both data-integrity verification and secure credential storage — two of the most common real-world security controls I’ll need to evaluate or implement.
😄 TL;DR
A hash isn’t a weaker lock — it’s not a lock at all. It can’t be decrypted because nothing was ever meant to come back out, which is exactly what makes it good for proving nothing was tampered with.
