I was just thinking this morning that Intel segments their market by the class of person that is their customer.
"You're a consumer, you don't need error correcting memory! Only the upper class, I mean... your superiors... err.. sorry, enterprise customers need that feature. It's reserved for them, it is not for the likes of you."
Meanwhile AMD basically sells silicon by unit area. It's like buying gelato at the ice cream shop. You can ask for one scoop, two scoops, or three. You get to decide how hungry you are, that's it. There isn't a flavour with broken glass in it[1] served only to the working class stiffs, and with the glass-free gelato reserved only for the gentry.
[1] This is pretty much what a CPU without ECC is. It randomly crashes and corrupts your data. Look. Not every bite of ice cream has glass in it! There's very little glass in the ice cream. It's super rare, and Intel Pty Ltd tells you that this is an acceptable amount of risk for you to take, because you aren't as important as other people.
> This is pretty much what a CPU without ECC is. It randomly crashes and corrupts your data.
I use ECC RAM in systems with error counters.
Error correction is an extremely rare event, if it happens at all. This idea that non-ECC computers are crashing all the time due to memory errors isn’t true. I've also had unstable systems with ECC RAM and no error correction events, so it's not a magic bullet.
You can run something like memtest86 for days on end. You shouldn’t see any errors at all, unless your memory is bad.
The reality is that the average consumer doesn’t really want to pay the extra amount for ECC RAM. As you said, ECC has been available on AMD consumer platforms for a long time and there’s barely any uptake outside of people building servers and workstations with consumer-grade AMD CPUs.
The AMD (Consumer) ECC story isn’t perfect, either. It’s not officially supported so there’s a lot of debate around whether or not it’s actually working on certain consumer motherboards. It’s definitely not equivalent to the official ECC support of their higher-end parts.
Intel, on the other hand, actually did roll out a lot of i3 processors with full, official ECC support. These are (or were) a favorite for low-cost servers for this reason. It worked well and the support was great.
Error Correction isn't the important part of ECC: it's Error Checking. There was a very long thread recently about this in the context of row hammer over at the forums on Real World Tech. While ECC can't fix row hammer, it can let you know that something is wrong and the system operator needs to intervene. The same thing is needed on desktops. Without ECC you won't know that your system is failing and silently corrupting your data. Especially when DRAM vendors are shipping products that don't reliably meet the defined operating parameters in their data sheets. "Trust, but verify".
The extra cost for ECC is basically noise in the cost of a desktop system. I buy unbuffered ECC DIMMs for my AMD systems and they're basically maybe 5-10% more than a non-ECC DIMM. The SSD premium is more than that. Registered DIMMs have a higher premium mostly because they need to have additional ICs on the DIMM to handle the higher memory capacities that servers tend to support.
I really wish CPU vendors would just flat out require ECC on every system shipped. Knowing that main memory is failing would save so much time speaking as someone who has encountered DIMM failures in otherwise stable systems. Days of debugging what turns out to be a hardware issue is no fun.
Same thought from a slightly different perspective. Decades ago I built a networking stack from scratch. It's wasn't wonderfully reliable. Years later I figured out why. I was wet behind the ears back then, and I put my faith in ARCnet's 16 bit CRC. Had I put aside that blind faith for just a few seconds and looked at how many packets per second that we being sent versus the number that would get through then - well I didn't.
I'm going to claim it wasn't entirely my fault. ARCnet did automatic resends for you so if fixed what the CRC detected, and silently what it didn't detect through. I had no idea the errors were happening.
ECC is nice, but as you say I don't need fancy hardware to correct a malfunctioning RAM chip. I can do that perfectly well with a hammer. What I need is the hardware to tell when and where to deploy the hammer. Things were actually better 30 years ago - back then PC RAM had parity. It's cheap, fast and effective enough. I have no idea what went wrong.
> Without ECC you won't know that your system is failing and silently corrupting your data.
The likelihood that a RAM error is corrupting data silently is slim, given the same corruption could affect the RAM storing the OS kernel or your application stack.
If you have bad non-ECC RAM, you will notice issues that warrant further investigation (i.e. memtest). But that assumes you're savvy enough to understand what failing memory looks like.
It's not about "likelihood", it's about being able to trust the hardware. The CPU running in your system probably has ECC protected caches, and many RAS features that happen to be designed in because most CPU vendors design the features into their cores that are shared between consumer gear and enterprise servers. These features are a necessary part of building hardware that can be trusted with atomic scale features.
Furthermore, your comment reads like someone that has never spent weeks debugging issues that only happen under extreme load stress tests. A single marginal bit is not so simple to track down when memory allocation patterns are non-deterministic. I have no desire to go through that again if it can be avoided by spending an extra $20 per DIMM. My time is worth more than that. Hardware is not perfect, and anyone who trusts hardware blindly should really take a look at what's going on behind the curtain.
Fwiw, the nature of DRAM errors is different now compared to the studies published decades ago when cosmic rays would only disturb a single cell. These days minute physical manufacturing defects and electrical disturbances are the dominant failure modes. Educate yourself: read the papers published about row hammer. Hammering a row with reads can flip bits in adjacent rows, and even DRAMs that have supposed mitigation features can suffer from disturbances a row or two further away. It has already been proven that the DRAM vendors' hardware cannot be trusted! What more proof do you need that this is a real world requirement in 2021?
The real pain is you have no way to sort out the reason of crash. Without ecc, you don't know whether it is the driver you installed last day cause the crash, or the memory itself. Or even driver installed yesterday used the broken region of the memory. Who knows? And memtest won't even come close to 100% reproduce the problem. If it happens continually, then it probably work. For for something that probably crash 1 in a month? No, unless you are lucky enough. And that happen all the time if you overclock. That is basically why ddrm5 mandate in memory ecc. Because in that high frequency, data corruption is more likely to happen.
The most painful case I had was actually bad memory in a FIFO in an ethernet switch. A few months before one of the Red Hat releases, IT had installed a brand new shiny enterpricey Ethernet switch with lots of line cards and many ports. This was one of the early VLAN capable switches. During the release validation cycle, QA found that one of our stress test workloads was occasionally failing with single bit flips in filesystem data from time to time.
Being kernel developers, we were majorly concerned. There were lots of changes to filesystem, VM and other parts of the kernel, and we couldn't rule out that we had a bug where code was using a stray pointer or some other problem. The stress tests had successfully passed last release...
Ultimately we tracked it down to network data being corrupted. The shiny new ethernet switch was rarely flipping a bit in packets, then happily fixing up the packet's CRC and IP checksum so that the computers on either end of the network link had no idea that the data was mangled in the network fabric. Oh the hours of brain wracking pain that caused!
> The shiny new ethernet switch was rarely flipping a bit in packets, then happily fixing up the packet's CRC and IP checksum so that the computers on either end of the network link had no idea that the data was mangled in the network fabric.
> I buy unbuffered ECC DIMMs for my AMD systems and they're basically maybe 5-10% more than a non-ECC DIMM
Where? Are you talking DDR4-3200? These are 40-50% higher than non-ECC UDIMMs. It usually gets worse as capacity increases. Additionally, there are only one or two models of 1x32GB UDIMM in production. These sticks are quite rare. You don't just go to Newegg or Amazon unless you want to pay out the arse for them. And in addition to that, ECC has worse timings which are important for Ryzen.
> Registered DIMMs have a higher premium
You have this backwards. RDIMM are common and cheaper because that's what servers generally use. UDIMM ECC cost more because, well, they barely exist as a thing.
UDIMMs don't have worse timings.
The only bad thing is that they don't offer binned dies with XMP profiles on their SPD.
My Kingston Server Premier 32G sticks with Micron Rev. E 16Gbit dies (KSM32ES8/16ME) run these timings very reliably at 1.4V:
OC'ing ECC is generally a terrible idea, if that's what you're doing. Also, no one on Ryzen is considering CAS CL22 timings "good" on DDR4-3200 sticks. I think you need to go do research on non-ECC RAM. It really feels like everyone on this thread is so far outside the normal PC market for RAM that they are talking out of their ass here.
> OC'ing ECC is generally a terrible idea, if that's what you're doing.
How so? For example I can see when the higher summer temperatures becomes an issue and reduce my timings, then tighten them again as it gets colder. No more random crashes where I wonder if it's my RAM or something else.
Currently doing 3333MT/s CL14 on 4x16GB of 2666MT/s ECC. I use CL16 in the summer. Since the memory sticks barely are affected by the heat, and that ECC errors occur on the same slot of the motherboard regardless of which stick is in there, I assume its the memory controller on the CPU in combination with four sticks of memory that's the limiting factor here.
I even got my 2 32G sticks in the same channel to 3600, but it wouldn't train with antyhing better than CL28 or therabouts.
I heard that's likely due to the imbalance and/or quad-rank making the memory controller unhappy, as the PHY should be agnostic to the CL setting.
I mostly just did that test to check if I'd have to give up the 3600 speed if I'd upgrade from 64GB 128GB on this 5950X; I'm well aware that one shouldn't be running it single-channel, especially not if that would mean quad-rank operation.
The notably part is that I operate it at above 1.2V, which increases power consumption in a way that enterprise generally doesn't deem worthwhile.
I run DDR4-3600 CL18.
I do question though how OC'ing non-ECC is a better idea than OC'ing ECC.
Also, just for the record, I based my timings off of an XMP profile for a non-ECC version of the same die/capacity configuration (32G 2Rx8 w/ Micron Rev.E 16Gbit).
I also haven't seen 32G sticks that do better than CL22 at 3200 with just 1.2V.
DDR4-3200 is already fairly slow for a new system -- DDR4-3600 f.e. aligns with the 1800 MHz Fabric Clock in Ryzen 3 and makes a not insubstantial difference in performance, and (non-ECC-wise) is not expensive.
I have a feeling the cost is not just a money-wise 5-10%.
Amazon currently lists the 32GB Nemix DDR4-2666 UDIMM for $169.99 while a non-ECC UDIMM is $129.99. So you're right, at $40 it is closer to a 30% premium right now. But that one time cost of $40 per DIMM is worth it if it even saves me 10-15 minutes of time debugging. And I'm willing to pay that small cost adder on every new machine I get, because I have the battle scars that have pro It's not like I'm dealing with hundreds of systems, this is across maybe 2 or 3 dozen machines over the past 30 years.ven it's a worthwhile tradeoff.
The battle scars are as follows:
- The 2MB expansion card on my Amiga 500 turned out to have a couple of bad bits, and because I was a green techie back then, I just attributed it to general software instability. Years later a RAM test showed there was a problem. Sadly there were no memory test utilities shipped with Amigas back then.
- My first 486 system had a bad bank in the SRAM on the motherboard. Took years to figure out as it only really hit when heavy DMA occurred from a VLB SCSI card while the CPU was heavily loaded. Swapped the SRAM out and it was stable for a few more years of service.
- One of my K6 systems had a DIMM go bad. At least it failed badly enough that the system couldn't boot.
- I've had an Intel Xeon in a colocated server start throwing ECC errors in the L3 cache. Replaced and RMAed the CPU before it caused an outage.
- One system started throwing ECC errors because the power supply was marginal.
Now, please, show me how the extra cost of ECC is worthless to you when things like this happen in the Real World. Is your time spent debugging hardware failures really not worth the cost of ECC?
You have to be fucking kidding me. On a desktop Ryzen?? Most desktop PC builders are gamers. That's just a fact. They drive the market. That's why every motherboard looks like a gamery stealth bomber. Outside of Asrock Rack and one or two Asus Pro boards, there is simply nothing out there that matches what the server world is seeing with SuperMicro.
I have 1x32GB DDR4-3200 UDIMM ECC on my Ryzen. But I have no delusions about the cost. It was expensive as fuck compared to non-ECC RAM.
(And there's no way to get ECC memory with decent timings because DDR4 is binned to death, so even if you found some hypothetical Samsung B-die ECC SKUs they'd be like 2333 or 2666 MHz and you'll never get them to work at 3200 or 3600 with timings better than 80-80-80-10000 or so. Meanwhile 3200CL14 or 3600CL16 is pretty usual without ECC.
edit: On a second look, Mushkin will actually sell you 3200 CL14 ECC memory... 14-18-18-38 that is. At almost 10 quid on the gigabyte. Meanwhile you can get real CL14 non-ECC memory at around half that.
But these weren't just cosmic-ray-induced errors. The DIMMs with errors were far more likely to have more errors in the future, pointing to faulty memory cells.
"Across the entire fleet, 1.3% of machines are affected by uncorrectable errors per year, with some platforms seeing as many as 2-4% affected" is enough to worry me.
I prefer more ECC than less, but as a friend of mine noted, most of the data are JPEGs (or video for that matter) and bit flips there usually go unnoticed because they simply adjust the pixel color slightly. So the chance that 1-in-2048-million bits is actually an important metadata bit that corrupts data structures is pretty low.
As someone who has been aware of ECC RAM for a long time could you briefly explain a concrete case where that is useful? Like, in all my time using computers I don't think I've yet encountered an error caused by RAM, or.. Did I?
You probably didn't even realise it was a memory error. A lot of the instability of Windows can be attributed to the use of non-ECC memory in typical home PCs. Linux tends to run on servers with ECC memory, so at least some of the perspective of it being "more stable" comes from this anecdotal evidence.
The parent comment saying that they never see ECC errors in the wild is missing a few things:
- Server memory tends to be clocked lower than consumer memory, so errors are less frequent to begin with.
- The errors are not evenly distributed. Some memory sticks have a high error rate, others are virtually zero. There's batch-to-batch variations.
- I've done my own tests on hundreds of servers. We run burn-in tests for about 24-48 hours. About 95% have zero errors of any kind, but 5% have a high enough rate that putting them into production would be a mistake. ECC allows us to catch those bit errors instead of silently accepting them and allowing data corruption to creep in.
- I've personally had 3 different personal computers experience high memory error rates, to the point of multiple BSODs per day and data corruption. They all started off "good" and slowly turned "bad". The only reason I knew to look for memory corruption as the root cause is because of my extensive industry experience. A grandma using the same PC would have just blamed Windows for being unstable.
- Vendors like Microsoft simply ignore all crash error reports sent back by telemetry with only 1 or 2 samples, because those are virtually guaranteed to be caused by memory corruption, not programmer error. I've done similar memory dump collection and found that easily 30% of all crashes were unique in this way, suggesting that ECC memory could improve PC stability significantly.
- Suggesting that ECC memory is not needed because "good" memory doesn't need it and only "bad" memory is a problem is missing the point. All memory is bad, it's just that the bit error rates are different!
I recently had an interesting experience with bad non-ECC RAM. The only effect I saw for many months was a frequently crashing web browser. I was very annoyed and downgraded to the firefox LTS release to try to solve the issue. The problem went away for a while, then came back. Seemingly erratic the problem could resurface only to disappear a while later.
Eventually I got a corruption in a large Git repository with photos. First I suspected a disk error but reruns of "git fsck" reported different bad objects, run to run. How odd, so I ran memtest86 and it reported a bad 32 MB memory region at offset 4 GB. Never saw any kernel issues or other instability. Booting up the computer does not use up 4 GB of memory (Linux) but starting a bloated web browser does.
The computer is stable after mapping away the bad memory region by using GRUB_BADRAM. That was an new takeaway for me, a machine with ECC memory can do this bad blocks mapping automatically. It's not just beneficial in correcting single bit errors.
I would love to have ECC memory in my machine. I used it for my previous build in 2012 but I think the situation has gotten worse since then. Bigger price difference and ECC memory is not even available at the same clocks as non-ECC memory.
Practically speaking, ECC is extra insurance against faulty memory.
It's true that ECC will catch errors caused by cosmic rays, but in practice most memory errors are just from faulty memory. Large studies in Google datacenters showed that the errors were heavily concentrated in a small number of DIMMs: http://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf
The catch is that you don't really know if you have one of these faulty memory sticks. I always run memtest86 overnight to check for obviously faulty memory, but some of these errors could take months to manifest.
I had a laptop with a RAM cell that failed. It would show up in memtest86 in a matter of minutes, yet surprisingly I didn't notice the issue for a very long time in day to day usage. I always wonder if there were random bit flips in anything I worked with during that period, but I'll never know.
With ECC, it's just one less thing to worry about.
One part of the equation is that a shit-ton amount of memory modules are, in fact, overclocked (e.g. do you think your DDR4-4800 runs at nominal speed?). That surely contributes to memory being faulty.
DDR4 Memory these day are rarely runs at their default frequency (2666,2333) unless you buy a really cheap one. Most 3200/3600 memory are in fact overclocked(Just the factory pre-configure the overclock setting with xmp) except for a few ones that are natively 3200(some memory from micron are natively 3200 at 1.2v).
Well there's a bunch of problems. Could be high energy particles flipping bits, but that's very rare. Could be a row hammer attack flipping bits. Could be a flaw in the memory chip, dimm, dimm socket, motherboard traces, CPU socket, or inside the CPU.
Without ECC any of the above will crash an app, or even the entire machine. With ECC you'll get a log message, and if it's a single bit error it will be automatically repaired. If it's not fixable and it's in userspace that application is killed (and the kernel logs the error), if it's in kernel space the kernel will panic.
So generally it makes your machine more reliable, and easier to debug. Generally repeated errors are a fault of some kind and you can easily tell which dimm it is. Without ECC you end up troubleshooting all causes of crashes, and even if you know it's memory, you can't be sure which dimm it is.
So I think it's well worth the minimal premium to make your machine more reliable, and if it's unreliable it's much easier to track and fix the problem.
Amusingly without ECC heavy CPU use with parallel gcc compiles causes a particular error. Not sure if it's in the FAQ, but it's well know that the particular error means you have a CPU/RAM problem, common with memory errors, CPUs that are too hot, or overclocked CPUs.
Anecdote: I once lost a lot of files over a period of ~two years on a desktop PC running ZFS with bad non-ECC RAM. I figured out what was going on after a bunch of audible corrupt blips kept appearing in all my favorite FLAC rips as they got mapped into the bad memory, checksummed against the contents still on disk, and written back "corrected" as far as ZFS was concerned. There's a better (i.e. not by me) write-up of it in the OP of this thread under "What happens when non-ECC RAM goes bad in a ZFS system?": https://www.truenas.com/community/threads/ecc-vs-non-ecc-ram...
Thank you, this is a perfect illustration and one of the scenarios I feared around keeping data long-term, though I was thinking about disk errors and such.
Sidenote I seem to have experienced Cunningham's Law for the first time semi-consciously, as I was aware my claim is most probably wrong. Still, I wonder if data corruption is a more common scenario rather than crashing, because ever since I use Linux I don't recall experiencing many crashes I couldn't attribute to a driver issue or known instability in the software I used, especially on compatible hardware. Windows days were another story... But maybe the memory modules were worse back then and perhaps having less memory made it more likely to crash?
How would you ever know? If your computer or a program has ever crashed or acted weird. That's a symptom of memory errors. It's also a symptom of a million more common issues but you would never know.
There are some bugs that get raised on linux once every few years that are so obscure they are suspected to be random memory errors.
You will find it useful if you crashed 1 in a month and realize turn the memory frequency down for a bit fixed it forever.(Technically I am still guessing, without ecc, you can't get the answer) Seriously, the hardware should be able to determine whether the configuration is safe or not instead of by randomly guessing yourself. Even the memory labeled as work as 3200, it isn't necessary the case when you actually use it.
It has nothing to do with the "class of person" -- that whole notion seems pejorative in this context -- but maximizing how much you can extract in the aggregate.
In the case of Intel, for years (decades?) they were their own primary competitor. They wanted to be sure that for a given customer they could extract the maximum amount possible, and they did this by gating some features (ECC, AVX, etc) to try to avoid business customers deciding to get by on "lesser" processors. And the simple truth is that the overwhelming bulk of consumer products will never, ever have an ECC relevant error event, which is how businesses managed to get by without it.
Slightly different in the AMD case as the Ryzen Pro are the same socket and not really much more expensive (eg, https://www.amazon.com/Ryzen-4750G-Processor-3-6Ghz-Threads/... ). So as a consumer you can make a fairly straightforward value prop tradeoff there. Well, you could except AMD throws in the complication that they no longer sell APUs directly to consumers, not officially anyway. These are all technically "OEM only" parts.
Official support, doesn't mean it doesn't work. Although your right AMD does a bit of product segmentation here/there. The laptop space while considerably simpler is not only segmented by core count, but threading, and TDP.
I've been expecting them to just simplify to a core count/memory channel/socket matrix for a while because they are so close. There really isn't any reason these days to segment by frequency beyond maybe a couple golden parts where all the cores run at the peak frequency.
Do you think every price should be set according to marginal cost? Since the marginal cost of a chip is extremely low after you've paid for the fab and design costs, should those be free as well?
In general you want to get more money from people who have more money. Everyone does this, including nonprofits and governments. There are lots of ways to wrap it up in things like value, status, or features and benefits, but ultimately that’s the basic goal.
A long while ago, I worked on part of a memory system that during development we knew was corrupting RAM. The amazing thing, is just how high the error rates could get before the machine would actually crash. That is largely because so much of RAM is code that doesn't get executed and data structures that are resistant to random corruption. Even actual dat corruption when it happens is frequently subtle because it flips a transaction timestamp by one digit in the midst of millions of transactions/etc, or the data is transient a video frame with a miscolored pixel, etc. Its really like playing Russian roulette with a revolver that has a few (b/m)illion chambers.
Generally by the time you see actual program crashes or notice data corruption your system is really good and screwed. That is part of why people in the know are so afraid of RAM corruption. It can be persisted and exist silently for a very long time before someone tries reading some file/transaction that was corrupted during a write years back, or the database suddenly starts crashing after it updates some index as part of a GC pass/whatever, while the actual bug/HW failure will never be reproduced.
If you luck out with reliable hardware the rate of pure bit flips from high energy particles is rather rare.
Problem is if you start having unreliable hardware it could be any component of your system. ECC memory helps track down a class of errors that could be in the ram chip, dimm, dimm socket, motherboard, CPU socket, or CPU.
My desktop cost another $100 or so because I got a E3-1230 xeon instead of the similar clocked i7. I bought it in 2015 and it's been fast, and reliable. Sure I might manage 6 month uptimes anyways, but I also might track down a dimm problem in hours instead of weeks.
Not that it happens frequently, but if it does and your application scenario is critical, how do one knows if say the corrupted data is a minor color component in an image or frame of a movie, or a variable in a spreadsheet cell buried somewhere? Would anyone notice? If a bitflip happened in a location containing the target address of a jump then the software or the OS would likely bomb immediately telling there's something wrong somewhere, but (unfortunately) there are subtle errors that can go unnoticed until it's too late. ECC is meant to protect also from those.
"You're a consumer, you don't need error correcting memory! Only the upper class, I mean... your superiors... err.. sorry, enterprise customers need that feature. It's reserved for them, it is not for the likes of you."
Meanwhile AMD basically sells silicon by unit area. It's like buying gelato at the ice cream shop. You can ask for one scoop, two scoops, or three. You get to decide how hungry you are, that's it. There isn't a flavour with broken glass in it[1] served only to the working class stiffs, and with the glass-free gelato reserved only for the gentry.
[1] This is pretty much what a CPU without ECC is. It randomly crashes and corrupts your data. Look. Not every bite of ice cream has glass in it! There's very little glass in the ice cream. It's super rare, and Intel Pty Ltd tells you that this is an acceptable amount of risk for you to take, because you aren't as important as other people.