Geekbench 5 is famously much better than Geekbench 6 for benchmarking these workstation-class CPUs (though many other benchmarks are better still, especially if you can grab the source code and compile them yourself).
I still like having a set of benchmarks that run across Android, iOS, Windows, macOS, Linux, and on Arm, X86, RISC-V, etc... even if imperfect, it's a point of reference to get a general feel. And the single core scores are a great representation of a 'feel' against baseline in day-to-day use.
Benchmarks like 7Zip compression/decompression, LAME encoding, Blackmagic RAW video encoding, Cinebench, x265 encoding, Blender encoding, Y Cruncher, and any of the built-in video game benchmarks make much more sense than Geekbench.
Some good benchmark options here https://hwbot.org/benchmarks
https://www.pcgamingwiki.com/wiki/List_of_games_with_built-i...
If you ever want a cheeky laugh ask your llm of choice to write a satirical Userbenchmark amd review.
Yes, they are imperfect, but they do broadly measure how fast a processor is, and can be used for comparison.
Benchmarks like 7Zip compression/decompression, LAME encoding, Blackmagic RAW video encoding, Cinebench, x265 encoding, Blender encoding, Y Cruncher, and any of the built-in video game benchmarks make much more sense than Geekbench.
I don't think so. First, Geekbench shows sub scores which is similar to those applications you mentioned. Second, nearly no one uses CPU rendering for Cinebench and Blender. They're mostly GPU work. Video encoding/decoding work is mostly done by the media engine or GPU on Apple Silicon. Third, CPU benchmarks tend to correlate. If one CPU is faster in one thing, it's more likely to be faster in another. Therefore, a comprehensive score like what Geekbench and SPEC provide is valuable. If you want specific applications, look at the Geekbench subscores.Geekbench is extremely good at showing general CPU performance, especially ST. It's also highly correlated with SPEC at nearly 1:1 in terms of scores as shown by Nuvia before they were purchased by Qualcomm.[0]
[0]https://medium.com/silicon-reimagined/performance-delivered-...
The Geekbench 5 approach of pretending Amdahl's Law doesn't exist is sometimes a valid benchmarking strategy, but generally is the wrong choice for benchmarking consumer workloads and devices, and that's what Geekbench is ostensibly targeting.
The fact that Geekbench 6 scores don't increase linearly with the addition of more CPU cores is not a weakness of the benchmark, it's the benchmark demonstrating an important real-world effect.
The change that Geekbench 7 makes to exclude some subtests from the multicore suite entirely will definitely have the effect of making the overall multicore score scale better with the addition of more cores, but most of the audience for those scores is going to miss out on the fact that the multicore test now measures a narrower range of tasks than the single-core test suite.
Why are they incomparable? If I want to run a highly parallel task surely this tells me which one to use?
> Benchmarks like 7Zip compression/decompression, LAME encoding, Blackmagic RAW video encoding, Cinebench, x265 encoding, Blender encoding, Y Cruncher, and any of the built-in video game benchmarks make much more sense than Geekbench.
Pretty sure the internal tests it uses are similar to these
You can extract meaning from them! Find your current computer, find the target computer, calculate the percentage difference. 1000 is 1% better than 990, so it indeed doesn't mean anything at all.
.. but it chugs power like it's not even funny.
I do like synthetic benchmarks as one input signal when evaluating a CPU; otherwise we go back to stupid numbers like frequency... which also never made sense because they had 1.8GHz AMD Athlon CPUs outperforming 2.4GHz Pentium 4's in 2002-3...
Doing a lot of real-world application testing is good, but there's always some you miss, and the worst part about it is that people start to lose patience when assessing.. How is a cinebench comparing to doing gamedev in UE5?
How is doing gamedev in UE5 comparing to doing gamedev in Snowdrop?
How does doing gamedev in Snowdrop compare to doing virtualisation (different CPUs handling that particular task better than others).
there's so many dimensions that it will always be true that there's not enough testing.
I'm not saying we shouldn't have rigorous testing like you say, in fact, what I'm actually saying is that "single number→would you like to go deeper" is a better pipeline than an excel spreadsheet that goes on for 10 pages (which still skips a bunch of nuance) and still better than one that goes on for 500 pages with a broad spectrum of topics.
The reason I'm so bitter about this is because outlets like LTT and Gamers Nexus seem to choose their games randomly (based on popularity I guess?) and so all the titles I worked on kinda got smeared by my "company"- but our games used hugely different game engines that work differently (Dunia vs Snowdrop handle CPUs... VERY differently. Snowdrop is built for multi-core. Anvil works best on a single core, and Dunia has trouble going passed 4 cores- so on large systems with multiple NUMA zones the performance is.. worse.
So the "battery of tests" always gave us a shitty score and reviewers just moved on believing themselves to be comprehensive arbiters of truth regarding performance and giving verdicts regarding the performance of games as a broad topic (which people take and don’t dig deeper themselves) and without regard to the utility of what they’re saying (500+ fps on CSGO isn't even renderable for example).
People might have been better served by looking broadly at how powerful the CPUs actually are and then digging in for their use case once they whittle a few down, because then also: developers will actually make software that tries to use the features on offer instead of assuming people will just buy CPUs that work better on their workloads.
It's a kind of "to big to fail" mentality where incumbents start being able to direct CPU sales based on those CPUs being optimised for their workloads. It's self-reinforcing.
It focuses more on ML than LLMs though.
> In Geekbench 7, a workload only runs in multi-threaded mode if the task it models actually runs multi-threaded in real applications. For example, the HTML5 Browser test isn’t included in the multi-threaded suite because web browsers are single-threaded (or lightly threaded).
But, of all of the 2-pick-only options, choosing a custom mix and calling it multi-core requires guessing your actual workload on both the content of the mixed test suite as well as the scaling profile of the CPU to reason with. On the other hand, just testing "single core" and "all core" at least only requires you to guess based on what you think the scaling profile of the CPU is.
Nothing beats just testing your actual workload, but that doesn't mean all other ways of testing have to be equally good.
For instance, the file compression workload on version 6:
This workload compresses and decompresses the Ruby 3.1.2 source archive (a 75 MB archive with 9,841 files) using different compression codecs (such as LZ4 and ZSTD). It also verifies the files using the SHA1 (Secure Hash Algorithm 1) function.
Here's research done by Nuvia before they were bought by Qualcomm: https://medium.com/silicon-reimagined/performance-delivered-...
(One could just as easily argue that the single-core test suite ought to exclude any task typically done with multiple threads in real-world usage.)
Yes, that's pretty much the whole point of Geekbench: to offer non-experts a way to measure CPU performance and produce scores that are relevant to a specific class of devices and users and use cases, rather than artificially inflated.
Benchmarking experts can already use something like SPEC, which does offer single-thread and both types of multi-thread test modes (SPECrate and SPECspeed).
CPU: https://www.geekbench.com/doc/geekbench7-cpu-workloads.pdf
GPU: https://www.geekbench.com/doc/geekbench7-gpu-workloads.pdf
Geekbench isn't a server CPU benchmark that stresses 128 core CPUs.
> The Clang workload uses the Clang compiler to compile the Lua interpreter, a popular open-source language interpreter.
People really shouldn't call Geekbench a synthetic benchmark anymore, since it has been using real world workloads for years now.
Geekbench 7, the latest version of Primate Labs’ cross-platform benchmark, has arrived and features both new and improved workloads to measure the performance of your CPUs and GPUs. Geekbench 7 is available for download today for Android, iOS, Windows, macOS, and Linux.

Geekbench 7 on Windows 11
Geekbench 7 includes new media workloads that measure how well your CPU handles audio and video encoding, decoding, and processing. The workloads model the tasks behind video conferencing, screen sharing, and everyday content consumption. These new workloads:
Alongside the new media workloads, Geekbench 7 adds a Game Physics workload built on the Jolt Physics engine used in popular video games, expands the Photo Editor workload with a richer set of real-world edits, and updates the Photo Library workload to support importing and processing modern image formats such as JPEG XL and DNG.
The multi-core benchmark in Geekbench 7 has been redesigned to better reflect how real applications behave. Not every task in the real world is multi-threaded, and pretending otherwise distorts scores without telling you anything useful about your device.
In Geekbench 7, a workload only runs in multi-threaded mode if the task it models actually runs multi-threaded in real applications. For example, the HTML5 Browser test isn’t included in the multi-threaded suite because web browsers are single-threaded (or lightly threaded).
The result is a multi-core score that’s a more accurate, more useful, and more relevant measure of how your device performs the work you actually do.
The GPU benchmark has a new focus on the machine learning and content creation applications that increasingly define GPU performance. For machine learning, the GPU benchmark includes workloads that:
Geekbench 7 also introduces new GPU image editing and synthesis workloads, including RAW image processing, LUT-based video color grading, path tracing, and fluid simulation.
And by popular demand, CUDA joins OpenCL, Vulkan, and Metal as a supported API in the GPU Benchmark. You can now measure your NVIDIA GPU using the API that powers its most demanding applications.
Since the release of Geekbench 6, the tasks people perform have become more strenuous, and the data sets they use have grown larger and more demanding. To reflect this, we’ve updated the data sets the workloads process so they’re more challenging for your device and better reflect the files people work with today. This includes:
Geekbench 7 is free (and will remain free) for personal use. Download it today and see how your devices measure up.
We’re celebrating the launch with 20% off Geekbench 7 Pro on the Primate Labs Store until August 6. Whether you’re a tech enthusiast researching your next upgrade or an IT professional managing a fleet of machines, there’s never been a better time to find out what your hardware can really do.