← Back to Overview
PUBLICATION TIMESTAMP
--

Geekbench 7 Rewrites the Rules of CPU Benchmarking—and Exposes Some Uncomfortable Truths

Geekbench 7 Rewrites the Rules of CPU Benchmarking—and Exposes Some Uncomfortable Truths

Last week, a curious thing happened in the quiet corners of Reddit and MacRumors: users started running Geekbench 7 on their brand-new 2026 flagship phones and watched the numbers crater. A Samsung Galaxy S26 Ultra with the latest Snapdragon 8 Elite Gen 5 managed decent scores. But an identically specced Xiaomi 17 Ultra posted single-core results less than half that of the Samsung. A OnePlus 15, using the same silicon, landed outside the top 10 altogether. And in a twist that would have been unthinkable a month ago, last year's Galaxy S25 family, running the older Snapdragon 8 Elite, swept four of the top five spots in the CPU rankings. The cause wasn't a bad batch of chips or a mysterious performance regression. It was Geekbench 7. Primate Labs officially released the seventh major version of its cross-platform benchmark on July 28, 2026. Available across Android, iOS, Windows, macOS, and Linux, the update is the first full overhaul since Geekbench 6 arrived in 2023. On the surface, it looks and feels almost identical. Underneath, the changes are tectonic. A fundamentally rearchitected multi-core methodology, native CUDA support for NVIDIA GPUs, a new AMD Ryzen 7 7700 baseline replacing the old Intel Core i7-12700, and an entirely new suite of AI and media workloads combine to make Geekbench 7 the most disruptive version the company has ever shipped. And it's already exposing uncomfortable truths about how smartphone manufacturers have been gaming the numbers.

For years, the multi-core section of Geekbench operated on a simple premise: every workload gets thrown at every core, whether the real-world application would ever do that or not. That approach rewarded processors with high core counts and aggressive scheduling, even if the actual user experience didn't scale accordingly. Geekbench 7 throws that playbook out. Now, a workload runs in multi-threaded mode only if the real-world application it's modeling genuinely behaves that way. The HTML5 Browser test, for instance, is gone from the multi-core suite entirely. Browsers are still overwhelmingly single-threaded or lightly threaded, so forcing them across 32 cores was synthetic theater. The File Compression workload, on the other hand, scales across cores because compression libraries do parallelize. Early testing from Hardwareluxx on a Ryzen 9 9950X3D showed File Compression scaling at 5.32x from single-core to multi-core under Geekbench 7, a sign of genuinely improved thread utilization for tasks that actually benefit. Primate Labs' logic is straightforward: a benchmark should reflect how people use their devices, not how benchmark engineers design for theoretical peaks. The resulting scores are no longer comparable to Geekbench 6. An Apple M4 iPad Pro that posted 3,719 single-core and 13,635 multi-core in Geekbench 6 returned 3,197 and 12,959 in Geekbench 7. The hardware didn't degrade; the measurement simply changed. Hacker News lit up over the philosophical implications. User wtallis pushed back against the notion that scaling linearly with core counts is inherently good: "GB5 ran N independent copies of a workload, pretending Amdahl's Law doesn't exist. GB6 runs a workload that requires splitting across cores with inter-thread coordination overhead. That scores don't increase perfectly with core count isn't a weakness—it's the benchmark demonstrating an important real-world effect." Yet even wtallis acknowledged a tension: by excluding certain workloads from the multi-core suite, "the multi-thread test now measures a narrower range of tasks than the single-thread test." Geekbench has long walked a tightrope between simplicity and accuracy. Version 7 doesn't solve that tension so much as it picks a different lane—one that prioritizes real-world workload behavior over theoretical throughput. For anyone who's ever wondered why a 32-core workstation felt no snappier in Lightroom than a 16-core machine, this lane makes intuitive sense.

The Phone Industry's Rude Awakening

That decision has collateral damage, and it's landing hard on smartphone manufacturers. The industry has spent years building elaborate scheduler optimizations that detect popular benchmarks and instantly crank frequencies, disable thermal throttling, and preload memory buffers. When Geekbench launches, the phone knows. It responds accordingly. Geekbench 7's redesigned workloads and execution patterns appear to have broken those detection mechanisms almost completely. The result is the score inversion that erupted on Reddit: phones from 2025 now outranking 2026 replacements, sometimes by wide margins. "The new run structure makes manufacturers' 'whitelist' mechanisms unable to recognize the benchmark," read one widely shared analysis. "Phones are reverting to their default power-saving behavior, releasing performance based on app identity rather than actual workload." Only vivo's iQOO 15 managed to buck the trend among Snapdragon 8 Elite Gen 5 devices, topping the charts while competitors from OnePlus, Honor, Xiaomi, and Samsung fell behind. That single exception points to a difference in scheduling logic rather than any hardware advantage, suggesting this isn't a chip problem but a software one. Meanwhile, a leaked Geekbench 7 run for the Galaxy Z Fold 8 showed a single-core score of 2,861 and multi-core of 7,840, numbers that will likely fuel debate about Samsung's tuning choices once more devices land in reviewers' hands. Japanese tech press noted that Sony's Xperia 1 VIII saw minimal score decline under Geekbench 7 compared to rivals from Xiaomi and OPPO, which suffered dramatic drops. In a single benchmark generation, Sony's traditionally conservative performance profile became a competitive advantage—not because its hardware got faster, but because its software wasn't tuned specifically for the previous version's quirks. Primate Labs didn't stumble into this. The company had already signaled its intentions with Geekbench 6.7 in April 2026, which introduced active detection of "BOT mode" and began invalidating scores flagged as artificially optimized. Geekbench 7 is the logical next step: make the playing field so different that the old cheat codes stop compiling.

New Workloads That Actually Feel 2026

Beyond the multi-core fight, Geekbench 7 adds an entirely new category of CPU tests built around media processing. There's a Video Encoder workload that takes screen-sharing footage and encodes it using the AV1 codec via the AOM library—directly modeling what happens when you share your screen on Zoom or Teams. An Audio Encoder compresses music and spoken-word tracks with Opus, mirroring voice memo and podcast apps. And a Video Decoder test decodes AV1 video and Opus audio while simultaneously generating live captions using OpenAI's Whisper speech recognition model. That last one is a particularly 2026 thing to benchmark. Live captioning, once a niche accessibility feature, is now table stakes for any video conferencing platform and a growing number of consumer video players. Measuring how well a device handles real-time AI transcription while also decoding media captures a workload that didn't exist when Geekbench 6 shipped. On the gaming side, a new Game Physics test uses the Jolt Physics engine—the same library powering collision detection in Death Stranding 2: On the Beach, Horizon Forbidden West, and War Thunder. This isn't a synthetic physics simulation; it's an actual game-industry physics engine, which gives the test real credibility with developers. One important note: Game Physics only contributes to the single-core score, so it won't inflate multi-threaded results artificially. Existing workloads got refreshed too. Photo Editor now handles a richer set of real-world edits, Photo Library supports JPEG XL and DNG, PDF Viewer renders six documents through PDFium (Chrome's rendering engine), and File Compression now runs three archive types across LZ4, zlib, and Zstandard with SHA1 verification. The datasets are larger, more varied, and more demanding than before—Primate Labs is deliberately raising the floor to keep the benchmark relevant as baseline hardware improves.

GPU Gets CUDA, Finally

For years, NVIDIA GPU owners running Geekbench were stuck with OpenCL or Vulkan—perfectly capable APIs, but not the ones that most professional and AI applications actually use. Geekbench 7 adds native CUDA support, letting NVIDIA hardware run workloads through the same API that powers the company's entire AI and professional computing stack. Digital Trends called it "perhaps the biggest addition," and the community reaction bears that out. On MacRumors, where Apple Silicon Metal scores have long dominated Geekbench's GPU charts, the arrival of CUDA means direct comparisons between platforms on workloads that matter to developers: AI inference, content creation, scientific computing. The GPU benchmark itself has shifted focus away from traditional 3D rendering toward machine learning and content creation. There's face tracking with real-time filter effects (think social media apps), AI-based image upscaling via RFDN (scaling a 256×256 tile to 1024×1024), background blurring for video conferencing using DeepLabV3+, RAW image processing with noise handling and demosaicing, LUT-based video color grading with tetrahedral interpolation, path tracing on the Blender BMW scene, and fluid simulations. To anchor all this, Geekbench 7 sets a new GPU baseline of 100,000 points, established by a Lenovo Legion laptop carrying a GeForce RTX 4060.

[SPONSORED]

▶ ENTERPRISE GPU CLUSTERS ◀

Scale your AI model training seamlessly. Book a Demo.

AMD Takes the CPU Baseline Crown

The CPU baseline has shifted from an Intel Core i7-12700 (a Dell Precision 3460) to an AMD Ryzen 7 7700 (a Lenovo Legion), normalized to 2,500 points. This is more than a cosmetic change. As Hardware Busters put it, the switch is "a clear marker of how much the mainstream desktop market has shifted since Geekbench 6 launched." Intel hasn't issued any public comment on the change, and that silence feels louder than any press release could be. Preliminary data from Primate Labs offers an early view of the new CPU hierarchy under Geekbench 7: | Processor | Cores | Single-Core | Multi-Core | |------------------------------|-------|-------------|------------| | Apple M5 | 10 | ~3,753 | ~18,722 | | Apple M4 Pro | 14 | ~3,360 | ~23,188 | | Qualcomm Snapdragon X2 Elite | 18 | ~3,225 | ~22,735 | | AMD Ryzen AI Max+ 395 | 16 | ~2,742 | ~27,412 | | Intel Core i9-13900KF | 24 | ~2,734 | ~24,439 | Apple's M5 leads in single-core, while AMD's Ryzen AI Max+ 395 takes the multi-core crown in this early sampling. The Snapdragon X2 Elite lands in a competitive position, though its scores under Geekbench 7 are notably lower relative to Apple silicon than what Geekbench 6 suggested. Qualcomm had previously highlighted a 4,033 single-core score in GB6 as evidence of leadership. In GB7, the narrative shifts.

The Benchmarking Industry Reacts—or Doesn't

UL Solutions, the company behind 3DMark, hasn't issued any public response to Geekbench 7's expansion into GPU territory. That quiet might be strategic. Geekbench's GPU tests emphasize machine learning and content creation, while 3DMark remains firmly focused on gaming. The two suites are now overlapping less than a casual glance might suggest. On Hacker News, the eternal debate over synthetic benchmarks re-erupted. One commenter dismissed them as "borderline useless—1000 is better than 990 but it literally doesn't mean anything." Others pushed back hard: "Geekbench is extremely good at showing general CPU performance, especially single-thread. It's also highly correlated with SPEC at nearly 1:1, as shown by Nuvia before Qualcomm acquired them." Academic research supports both sides: Geekbench users show a mean absolute error of 11.2% compared to 5.5% for SPEC users, but that gap reflects differences in user expertise and testing conditions, not necessarily test quality. Professional reviewers are still finding their footing. Gamers Nexus hasn't published its full analysis yet—reasonable, given the need to accumulate data across multiple hardware configurations. Hardware Busters provided the most detailed early coverage, noting that the launch itself was "chaotic," with Primate Labs providing embargo materials to outlets but then releasing everything publicly before anyone could publish properly. "Effectively, they self-lifted the embargo," the site observed.

What This Means for Buyers (and Marketers)

For anyone buying a laptop or phone based on benchmark scores, the transition is messy. Geekbench 7 scores are not comparable to Geekbench 6, and every historical comparison chart is now invalid. Consumer organizations and review sites will need to rebuild their performance databases from scratch. But Geekbench 7 also means manufacturers can no longer coast on artificially inflated numbers from benchmark-specific optimizations. The new AI workloads—face tracking, super-resolution, background blur—directly map to features that marketing departments love to highlight. A phone that does well in Geekbench 7's AI tests can credibly claim leadership in exactly the kinds of workloads that appear on spec sheets: live filters, on-device image enhancement, video call effects. And because Geekbench supports RISC-V alongside x86, Arm, and Apple Silicon, emerging architectures now have a seat at the comparison table, which matters for server and embedded markets alike. Primate Labs has kept the pricing model unchanged: free for personal use, $99 for Pro (with a 20% launch discount to $79 through August 6, 2026). Pro enables offline benchmarking and covers all desktop platforms. Some MacRumors forum users grumbled about having to repurchase for each major version—"Geekbench perpetually makes us buy it again," wrote one—but the price point hasn't shifted upward. Founder John Poole has yet to give a detailed interview about Geekbench 7's commercial strategy, though his history of surfacing uncomfortable truths (most famously, Apple's iPhone throttling controversy) suggests a certain comfort with industry disruption. Geekbench 7 continues that tradition, whether the smartphone industry loves it or not. The download is live now at geekbench.com. Early adopters are already flooding the Geekbench Browser with results, and the new performance landscape is taking shape in real time. Whether the benchmark's new realism will translate into better buying decisions—or just a different flavor of score chasing—is now up to the community of reviewers, manufacturers, and users to figure out.

Editorial Disclosure: This commercial analysis is compiled from global informational platforms and developer community discussions. Due to rapid technical cycles, readers are advised to independently verify volatile metrics. FUTUREMARSNEWS maintains structural objectivity and independent neutrality. more
This publication is intended solely for commercial, educational, and informational purposes. Articles may include news reporting, editorial opinions, technical analysis, software tutorials, deployment guidance, benchmark testing, hardware evaluations, workflow optimization strategies, pricing references, market intelligence, developer resources, and enterprise technology commentary. Product specifications, APIs, licensing models, cloud pricing, benchmark results, software capabilities, commercial terms, and hardware availability are subject to change without notice. Any performance figures or comparisons are based on publicly available information, vendor documentation, independent testing, or specific test environments and should not be interpreted as universally representative. Readers are encouraged to verify all technical and commercial information directly with official vendors before making engineering, purchasing, investment, or operational decisions. Unless explicitly labeled as sponsored content, advertising, affiliate content, or paid partnerships, editorial decisions remain independent. FUTUREMARSNEWS does not warrant the completeness, accuracy, or future availability of third-party products, services, software, or information referenced within this publication.