Aug 312020
 
Pentium II Overdrive logo by Lukas Wojdyla / lukparts

[1] 0. Introduction

It’s taken years of waiting before it could come to this: The final CPU upgrade the XIN.at server will ever have seen, as [vaguely announced]. Being an IBM PC Server 704 8650-4M0, it features four socket 8, supporting up two four Intel Pentium Pro 200MHz 1M CPUs as the absolute maximum. And to make the 1M CPUs from 1997 work, you need IBM 12J3352 CPU riser boards as upgrades as well.

Now there is one particularly interesting thing when it comes to the socket that hosted the first real 686, and that’s that there is an Overdrive chip for it. For those of you too young to know: Back in the 90’s, when upgrading a whole computer was seriously expensive, Intel offered special upgrade processors for older platforms, so you could e.g. plug a Pentium processor into a 486 mainboard. And the final Overdrive Intel ever made was for socket 8, so you could replace your Pentium Pro chip(s) with a Pentium II (actually, Xeon) class CPU: The Intel Pentium II Overdrive, which Intel themselves only officially supported for single and dual socket systems, but not for quad socket, like with the machine that is hosting this web site.

1. A bit of backstory & the first test

I always thought that the IBM server would most likely not allow for any CPU to run, unless it’s supported by Intel and IBM for that machine. Reason being that I needed new riser boards to get those 1M Pentium Pro’s working. Here’s one of my old chips, or rather a spare I had lying around, as usual, click to enlarge:

Intel Pentium Pro 200MHz 1M processor

Intel Pentium Pro 200MHz 1M processor

Now there, a German guy named [S2 Sedan]German flag – salvager on the side – has been looking for old IBM PC Server 704 machines for me for many years. Chances to find any was really low, but just recently, he hit the jackpot, securing a good source of spare parts for me. Just look at that, including the sacrilege that I ordered to be committed:

 

The rightmost picture still makes me wanna cry… But, to get back on topic, S2 Sedan also managed to get one of them to boot up, and given the nature of his side business, he was already in possession of four Pentium II Overdrive chips in their retail version:

S2 Sedan's Pentium II Overdrives, shown on ASUS C-P6ND riser boards

S2 Sedan’s Pentium II Overdrives, shown on ASUS C-P6ND riser boards [2]

So I had sent him two IBM 12J3352 spare riser boards I had in stock, and he tried to get it up and running. Some jumper configurations and some switching servers was necessary, but in the end:

 

I couldn’t believe my eyes! I mean, of course this didn’t mean that it would really boot up into an operating system and work in a stable fashion, but at the very least one can get past the power on self test, which is in itself already much more than what I’d hoped for!

Those photos really got me fired up, so I went to eBay USA and looked for Overdrives. I actually wanted the OEM versions without fans attached, and there is one seller who’s always selling a single OEM unit at any given time. I won’t link to it, but just search for “Pentium II Overdrive” on eBay.com, and you’ll find it. If you want multiple units for your system(s), just ask him, and he will likely be able to deliver. It appears he has quite a lot in stock, at least at the time of writing.

Thanks to Corona-chan and Pitney Bowes being a slow logistics company it took a month for them to arrive, but there they were!

2. The processors

The OEM versions have no flashy boxes, but are just sitting in sealed blister packages. Parts of the package rip / break when opening them, which was the case here. Those were truly still sealed?! Unbelievable.

A set of five Pentium II Overdrive 333MHz CPUs in their OEM versions

A set of five Pentium II Overdrive 333MHz CPUs in their OEM versions, four for using, one as a spare

More images:

 

The passive cooler is surprisingly small. Given the 45W TDP of the Pentium Pro 200MHz 1M chips, I was a bit concerned about this. I mean, the IBM PC Server 704 has thermal zones and really good airflow, but still. I couldn’t find any TDP specifications for the Overdrives, so I just looked at the regular Pentium II 333MHz with 0.25Âĩm “Deschutes” core, which is also a part of this upgrade CPU (the original “Klamath” Pentium II was fabricated at a 0.35Âĩm node). [According to CPU-World], it’s 23.7W, so not bad. A pretty cool chip in comparison, if we assume that the Overdrive would be roughly in the same ballpark.

The carrier board has some voltage regulators soldered onto it of course, so I tried to peek a little:

 

I couldn’t make out too much, so I decided to crack it open, also to inspect the state of the thermal grease on the CPU and its L2 cache chip:

 

I was really surprised to see that the black cover on the pins could be removed. We can see that the thermal grease is still in excellent shape, which was another surprise after 23 years in storage! Not bad. Aside from the core (the chip without heatspreader) and the larger cache chip, you can also make out some voltage regulators, as expected.

I reassembled the CPU and got ready for the first installation attempt.

3. Testing my new Eaton 9PX UPS unit

I haven’t actually written about this, but due to another not battery related UPS unit failure, I decided to give up on APC and switch to Eaton for my UPS units. I went a bit overboard and got the following to protect XIN.at recently (sorry, no hardware photos):

  • Eaton 9PX 1000W RT2U VFI (9PX1000IRT2U)
  • Eaton 9PX extended battery module (9PXEBM48RT2U)
  • Eaton Network Management Gigabit Network card (Network-M2)
  • Eaton Environmental Monitoring Probe Gen2 (EMPDT1H1C2)
  • Eaton Hotswap MBP (MBP3KID)
  • Eaton Cable kit for the Hotswap MBP for “small” UPS’s (CBLMBP10EU)

To be able to power off & reboot the machine remotely, I’m using a KVM-over-IP box with client software adapted & re-released by myself plus the power cycling capability of the UPS unit, controlled via its web interface. This had to be tested for the new Eaton 9PX as well, and this was the perfect opportunity. So I used my convertible tablet, went online fia LTE, and connected to my server using [XViewer], which is specific to a single TrendNet KVM box:

Accessing the IBM PC Server 704 via KVM-over-IP using XViewer

Accessing the IBM PC Server 704 via KVM-over-IP using XViewer

Here I would pick “Beenden” which means “Power off”. Then, since the machine has no ACPI yet, it would shut all processes down, flush all caches to disk and tell you that “You can now switch off the computer”. This is where the power cycle capability of the UPS comes in. It’s divided into two power groups, just like with the APC SmartUPS series. So you can power cycle one group (here: the server) while keeping a second one online (here: all the networking hardware):

Eaton 9PX output group power controls

Eaton 9PX output group power controls

Thankfully, everything worked perfectly, just as planned!

I had decided to take a highly compressed full backup of the server’s system drive (36GB 15000rpm SCSI) and the data RAID-5 (roughly 55GiB, also SCSI) afterwards. It took a whopping 12 hours to get it done on those sluggish CPUs, so it was around Saturday 04:00 am when the following pictures were made. ;)

4. Installation attempt

I opened the server and did some cleaning, of which I do not have any current photos, so I’ll just give you some really old ones so you roughly can see what it looks like inside:

 

I pulled the CPU riser boards from the machine, removed tha passively cooled Pentium Pro CPUs and plugged in the “new” Overdrives:

 

The interesting part here is that while the heatsinks are clamped into the carrier board – or rather the plastic cover underneath the bottom of it – there is no clamp included to fix them to the socket! So they’re fixed only by the friction between the pins and the socket. That’s a bit scary, given how two of the processors will be hanging heads down. I tested this outside of the machine for a night, and they didn’t fall out. I also grabbed a board by one of the heatsinks and lifted it up, tried to “shake” the board off the CPU. But that didn’t work, so it’s probably okay.

I’ll keep checking from time to time though…

After the CPUs were attached, I still needed to set the L2 configuration jumpers back from 1M to 512k for both sockets, so that the machine would boot up. Yes, there is actually a jumper for that, as specified in the [IBM Netfinity 7000 Hardware Maintenance Manual Supplement]:

IBM 12J3352 riser board, L2 cache jumper documentation

IBM 12J3352 riser board, L2 cache jumper documentation

Now my server isn’t an original Netfinity 7000, but the two are quite similar, and share some parts as well. Like the CPU risers.

So let’s switch them to 512k:

 

Given the opportunity, I also reinstalled the Delta WFB1212ME-R00 fan I [had serviced], and put the spare back into storage. More service time for that old thing!

When that was done, it was time to put everything back together, take a deep breath and push that button!

5. Power on

I was really happy to see the machine come back online! Here’s a comparison of P.O.S.T. screens during the CPU detection phase, on the left side with the old, officially supported CPUs and on the right side with the new Overdrives:

 

The system runs its FSB at 66.2MHz once the new chips are installed, resulting in a very mild underclocking from 333MHz to 331MHz, but that’s fine. It’s actually working! Naturally, lacking any Âĩcodes for those chips, the BIOS has no choice other than to show them as unknown processors or as Pentium Pro’s. Maybe the “Pentium Pro” string was even hard-coded into the BIOS, could be.

Now there’s a funny part too, look at this:

 

Uhm, alright, so the front LCD panel thinks I’m running 75MHz chips. That’d be quite the downgrade. ;) So what’s happening here? What I think: The LCD is attached to a controller board, which is in turn hooked up to the system board. I am assuming that the small memory area delivering the clock speed value to the LCD is likely just 8 bits wide. If true, then it can represent only 28 = 256 values. Since 0MHz make no sense, it’d likely start counting at 1.

That would mean that in this configuration the LCD would be able to support displaying speeds from 1..256MHz. But we’re running 331MHz. Here’s what supports my theory: 331 – 256 = 75!

This is a classic unsigned byte overflow! At 257, the first number it cannot represent, it just trips and represents that value as 1 again. Made me chuckle to see an overflow in hardware like this. ;) That’s what can happen when using unsupported processors.

The machine continued to boot happily after that, with the operating system coming up perfectly fine! Before the benchmarks, let’s look at another little detail!

6. Heat and power consumption

Like I said before, those Pentium Pro 1M chips were really hungry and hot. The die is massive after all, like the area of a grown man’s thumb, only wider. 45W of TDP vs. maybe around 27W? Let’s take an actual look:

 

Take a look at “Output group 1” only here, as group 2 drives the networking hardware only, and the first value denotes the power consumption of just the UPS unit by itself.

I already had a feeling when I put my hand behind the server, where the air from the CPU & RAM compartment is leaving the case. It was suspiciously cool, even under load. And you can clearly see why: With the CPUs changed, the whole system consumes 80 watts less power than before! That’s quite the amount! If we sum up the TDPs they’d amount to 180W for the Pentium Pro’s, and given our assumption from before is correct, roughly 95W for the Overdrives. Maybe 100W given the added voltage regulation circuitry.

And those numbers fit together really well…

So, more performance, less load on the CPU riser board voltage regulators, less heat, lower cost. Nice. Well, “lower cost” being a bit of an eyewash, given the amount I had to pay to get the CPUs. It’s not like they can amortize that so quickly. But hey, it’s okay, I’ve got time. ;)

7. Benchmarks

Note: All benchmarks were run while the server was in productive use (Web-, Mail-, IRC servers etc.), so results are to be taken with a grain of salt.

Now, I also did some quick (or not so quick, actually) benchmarks to show the difference between old and new. One test was to open two connections to my FTP server using the AES256-GCM-SHA374 cipher through the TLS v1.2 protocol, and to download two large files in parallel for some time while being limited to my maximum WAN bandwidth of 8Mbit/s.

Secondly, I benchmarked Cinebench R10 on it, to visualize performance scaling not just from Pentium Pro to Pentium II Overdrive, but also in terms of parallelism, from 1 CPU to 4 CPUs under load. Thirdly, I will run my own [x264 benchmark] in its special build for SSE-less CPUs on it, but that’ll probably take 2 weeks, so you’ll just have to sit and wait for that. Turns out it managed to complete it in less than a week. ;)

First, the FTP test:

 

Hmm, this looks rather inconclusive. We can see that the load is more compact across cores with the Overdrives, but I couldn’t make out a clear winner here. The load starts where the top traffic graph starts peaking in red. The reason why it’s not 2 CPUs being loaded constantly with 2 being idle is that the operating system’s thread scheduler is pushing the threads around all the time.

Even when trying to merge the graphs, it’s still not clear:

Load graphs merged

Load graphs merged, Pentium Pros in grey, Overdrives in color

Well, it’s mostly an unintelligible mess I guess. But even when looking very closely, one couldn’t determine any kind of clear result here. After ending the test on the Overdrives, I noticed that the base load of the server was just too high. You can also see that before the test begun; The seas were just calmer for the Pentium Pros at the time of the test. Maybe there was some load on the web server or something.

I will have to redo this on the Overdrives when the conditions are better for comparability. I will update the article later on to show my findings.

The thing is, several other services like the weblog interface (which is what I’m writing posts with) feel quite a bit more snappy, so I’m a bit disappointed that I couldn’t visualize this with FTP+TLS load. Well, we’ll see how it goes after a re-test, maybe tonight.

Update 2020-09-01:

Alright, I re-ran the FTP test, but unfortunately, the result is still somewhat inconclusive. Here’s the combined graph. I tried to make it more readable this time around:

Load graphs merged again

Load graphs merged again, Pentium Pros in green, Overdrives in magenta

Well, still hard to say. Maybe, uhm, a little bit better? I think I’ll change my test method, just to make sure. I’ll compare the speed over LAN instead of WAN, so I’ll fully load 2 CPUs, again with two parallel connections. Speed with the Pentium Pro chips was 1.3MiB/s per connection in that case, so 2.6MiB/s in total. Variation per connection was Âą0.2MiB/s. Once the x264 benchmark is through, I’ll check what kind of throughput I’ll get with encrypted FTPS on the local network.

End of update

Update 2020-09-07:

The encrypted FTP throughput test has been completed, and I think it speaks for itself. Even if we account for the variation, it’s been at most 3MiB/s in total before, and now?

FTP transfer test

FTP transfer test (click to enlarge)

While monitoring that transfer, the lowest I’ve seen was 1.9MiB/s per transfer, so 3.8MiB/s in total. The screenshot above shows the peak, 4.5MiB/s. In any case, it’s much better than before, so yeah!

End of update

Now, let’s take a look at Cinebench R10, which shows some really interesting results:

 

Now, here’s the thing: In pure single-CPU load, the performance increases by roughly +50%. That’s really not too bad given the clock speed increase is exactly +66%.

But it got more challenging when I started to load all of the CPUs. The MP ratio factor of the old Pentium Pro’s wasn’t all that shabby at 3.31×. Pretty healthy for an ancient quad socket machine such as this one. But with the Overdrives, that value drops sharply, down to 2.78×!

So when loading all CPUs with Cinebench, the performance improvement is only +25% when compared to the old processors. This might be another reason for Intel not supporting those chips in that configuration. They don’t scale well. And showing off diminishing performance increases when scaling up to 4 sockets might’ve not sold well at all. This kind of issue is likely going to be very application-dependent, but for processes doing lots of transfers from and to RAM, I’d think you’d see relatively bad numbers when utilizing all four processors.

My assumption is that the chips are just too fast for the underlying i450GX platform and its slow 66MHz FPM-DRAM memory subsystem. While the memory is 4-way bank interleaved to improve bandwidth, that interleaving doesn’t work all that well most of the time it seems. Probably needs large burst transfers to work well or something.

As for the x264 benchmark: As said it’ll take a while. It’ll be published here as soon as I have results ready! Of course including a comparison with the old CPUs. Hint: It’ll start pretty much now, so performance will remain impacted for at least two weeks starting with 2020-08-31!

Update 2020-09-07, x264 benchmark results are in:

My x264 benchmark has now concluded, and the results [can be seen here], in a direct comparison of the old vs. the new chips. I’ll give the results to you here as well:

  • 243:18:19.359 | 4 × Intel Pentium PRO 1MB 200MHz @ 199MHz | 2GiB ECC+P FPM-DRAM | IBM PC Server 704 8550-4M0 | Intel 450GX Orion | Windows 2000 Server SP4 (Custom GCC Build)
  • 159:08:01.156 | 4 × Intel Pentium II Overdrive 333MHz @ 331MHz | 2GiB ECC+P FPM-DRAM | IBM PC Server 704 8550-4M0 | Intel 450GX Orion | Windows 2000 Server SP4 (Custom GCC Build)

The result is better than I’d thought it’d be; For +66% clock speed we get +52.9% performance, which is quite respectable. This also supports my theory that processes which don’t use a lot of memory bandwidth (like x264) can scale pretty well, whereas processes which do will scale much worse.

End of update

8. Conclusion

 

Pentium Pro 200MHz 1M Pentium II Overdrive 333MHz 512k
Manufacturing node / transistor gate width: 0.35Âĩm (350nm) Manufacturing node / transistor gate width: 0.25Âĩm (250nm)
Core clock speed: 200MHz Core clock speed: 333MHz
L2 cache clock speed: 200MHz (full) L2 cache clock speed: 333MHz (full)
Level 1 instruction cache: 8kiB, 4-way set associative Level 1 instruction cache: 16kiB, 4-way set associative
Level 1 data cache: 8kiB, 2-way set associative Level 1 data cache: 16kiB, 4-way set associative
Level 2 cache: 1024kiB, 4-way set associative Level 2 cache: 512kiB, 4-way set associative
Instruction set extensions: None (pure 686) Instruction set extensions: MMX
   

A total downtime of 12 hours for the full backup and another 1-2 hours of impaired uptime due to CPU cycles being consumed by Cinebench. Was it worth it? Absolutely, in my opinion. While the results for full load don’t look all that perfect, and the FTP results remain inconclusive, I can tell from actually using the server that it’s faster now.

When I’m saying that, I mostly mean PHP performance in the web server context, which means single CPU load if the server is otherwise not overly loaded. It’s really noticable and “feels” more like what Cinebench’s showing. It’s still very slow by today’s standards of course, but given my workflow using the weblog software, it’s far more bearable now. Less time just sitting there, staring at the screen and waiting for the software to scale down that image or send that comment or edit that post. I still have to wait here and there, but I feel less handicapped and can publish things faster.

I’m thinking that when posting comments here now, you should be able to expect a 20-40% more responsive server. It’ll still suck of course, but at least a little less than before. ;) Also, it’ll suck for the 2 or so weeks that the x264 benchmark run will need to complete. ;)

Sending and receiving eMails also feels a slight bit more responsive, but of course, that only affects what few users are actually using XIN.at for their eMail services.

The IRC server’s M.o.t.D. / message of the day (A reminder: There is a [webchat interface] as well), [PRTG] and several informational pages here have already been edited to reflect the hardware change.

9. What else?

All that I’m waiting for now is the Pentium II Overdrive stickers I ordered from Lukas / lukparts, who also made my second batch of FreeBSD stickers, which I [used] on my AMD Threadripper box. He has shown me a prototype based on a design I sent to him already (the logo for this post), see the image below!

"Pentium II Overdrive Processor" case badge made by lukparts

As said, I will post pictures of the final product when it’s here!

I can’t wait! As soon as they’re here, I’ll post another picture of it on the case of my IBM PC Server 704! :)

As mentioned, updated FTP benchmarks, photos of that sticker when it’s arrived as well as other updates related to the use of Pentium II Overdrive CPUs in my IBM PC Server 704 will arrive as edits!

Update 2020-09-11:

And here it is!

Pentium II Overdrive case badge

Pentium II Overdrive case badge by [lukparts]

End of update

It took a really long time, but finally, the upgrade I’d thought most unlikely to work has been applied successfully! Beer Smilie

[1] Logo image is ÂĐ Lukas Wojdyla a.k.a. lukparts on etsy.

[2] Photographs are ÂĐ S2 Sedan.

Nov 192016
 
FreeBSD GMABoost logo

Recently, after finding out that the old Intel GMA950 profits greatly from added memory bandwidth (see [here]), I wondered if the overclocking mechanism applied by the Windows tool [here] had leaked into the public after all this time. The developer of said tool refused to open source the software even after it turning into abandonware – announced support for GMA X3100 and X4500 as well as MacOS X and Linux never came to be. Also, he did not say how he managed to overclock the GMA950 in the first place.

Some hackers disassembled the code of the GMABooster however, and found out that all that’s needed is a simple PCI register modification that you could probably apply by yourself on Microsoft Windows by using H.Oda!s’ [WPCREdit].

Tools for PCI register modification do exist on Linux and UNIX as well of course, so I wondered whether I could apply this knowledge on FreeBSD UNIX too. Of course, I’m a few years late to the party, because people have already solved this back in 2011! But just in case the scripts and commands disappear from the web, I wanted this to be documented here as well. First, let’s see whether we even have a GMA950 (of course I do, but still). It should be PCI device 0:0:2:0, you can use FreeBSDs’ own pciconf utility or the lspci command from Linux:

# lspci | grep "00:02.0"
00:02.0 VGA compatible controller: Intel Corporation Mobile 945GM/GMS, 943/940GML Express Integrated Graphics Controller (rev 03)
 
# pciconf -lv pci0:0:2:0
vgapci0@pci0:0:2:0:    class=0x030000 card=0x30aa103c chip=0x27a28086 rev=0x03 hdr=0x00
    vendor     = 'Intel Corporation'
    device     = 'Mobile 945GM/GMS, 943/940GML Express Integrated Graphics Controller'
    class      = display
    subclass   = VGA

Ok, to alter the GMA950s’ render clock speed (we are not going to touch it’s 2D “desktop” speed), we have to write certain values into some PCI registers of that chip at 0xF0hex and 0xF1hex. There are three different values regulating clockspeed. Since we’re going to use setpci, you’ll need to install the sysutils/pciutils package on your machine via # pkg install pciutils. I tried to do it with FreeBSDs’ native pciconf tool, but all I managed was to crash the machine a lot! Couldn’t get it solved that way (just me being too stupid I guess), so we’ll rely on a Linux tool for this. Here is my version of the script, which I call gmaboost.sh. I placed that in /usr/local/sbin/ for global execution:

  1. #!/bin/sh
  2.  
  3. case "$1" in
  4.   200) clockStep=34 ;;
  5.   250) clockStep=31 ;;
  6.   400) clockStep=33 ;;
  7.   *)
  8.     echo "Wrong or no argument specified! You need to specify a GMA clock speed!" >&2
  9.     echo "Usage: $0 [200|250|400]" >&2
  10.     exit 1
  11.   ;;
  12. esac
  13.  
  14. setpci -s 02.0 F0.B=00,60
  15. setpci -s 02.0 F0.B=$clockStep,05
  16.  
  17. echo "Clockspeed set to "$1"MHz"

Now you can do something like this: # gmaboost.sh 200 or # gmaboost.sh 400, etc. Interestingly, FreeBSDs’ i915_kms graphics driver seems to have set the 3D render clock speed of my GMA950 to 400MHz already, so there was nothing to be gained for me in terms of performance. I can still clock it down to conserve energy though. A quick performance comparison using a crappy custom-recorded ioquake3 demo shows the following results:

  • 200MHz: 30.6fps
  • 250MHz: 35.8fps
  • 400MHz: 42.6fps

Hardware was a Core 2 Duo T7600 and the GPU was making use of two DDR-II/667 4-4-4 memory modules in dual channel configuration. Resolution was 1400×1050 with quite a few changes in the Quake III configuration to achieve more performance, so your results won’t be comparable, even when running ioquake3 on identical hardware. I’d post my ~/.ioquake3/baseq3/q3config.cfg here, but in my stupidity I just managed to freaking wipe the file out. Now I have to redo all the tuning, pfh.

But in any case, this really works!

Unfortunately, it only applies to the GMA950. And I still wonder what it was that was so wrong with # pciconf -w -h pci0:0:2:0 0xF0 0060 && pciconf -w -h pci0:0:2:0 0xF0 3405 and the like. I tried a few combinations just in case my byte order was messed up or in case I really had to write single bytes instead of half-words, but either the change wouldn’t apply at all, or the machine would just lock up. Would be nice to do this with only BSD tools on actual FreeBSD UNIX, but I guess I’m just too stupid for pciconf

Jan 232014
 

Tulsa logoUpdate 2025-08-24: Tulsa has been benchmarked by now in the form of a quad Xeon MP 7140M system, [see the results]. If you look at the [full list] without filters applied (so you can compare the system to others), the faster custom build is #488 and the reference one is #684, thanks fly out to Biolante for running the benchmark on that insane system! :) End of update.

Over the past few years, my [x264 benchmark] has been honored to accept results from many an exotic system. Amongst these are some of the weirder x86 CPUs like a Transmeta Efficēon, a cacheless Intel Celeron that only exists in Asia, and even my good old 486 DX4-S/100 which needed almost nine months to complete what modern boxes do in 1-2 hours. Plus the more exotic ones like the VLIW architecture Intel ItaniumÂē or some ARM RISC chips, one of them sitting on a Raspberry Pi. Also, PowerPC, a MIPS-style chinese éū™čŠŊ, or Loongson-2f as we call it, and so on and so forth.

There is however one chip that we’ve been hunting for years now, and never got a hold of. The Intel TULSA. A behemoth, just like the [golden driller] standing in the city that gave the chip its name. Sure, the Pentium 4 / Netburst era wasn’t the best for Intel, and the architecture was the laughingstock of all AMD users of that time. Some of the cores weren’t actually that bad though, and Tulsa is a specifically mad piece of technology.

Tulisa Contostavlos

Tulisa? That you?

Ehm… I said Tulsa, not Tulisa, come on guys, stay focused here! A processor, silicon and stuff (not silicone, fellas).

Xeon 7140M "Tulsa"

An Intel Xeon 7140M “Tulsa” (photograph kindly provided by Thomsen-XE)

Now that’s more like it right there! People seem to agree that the first native x86 dual core was built by Intel and that it was the Core 2. Which is wrong. It wasn’t. It was a hilarious 150W TDP Netburst Monster weighing almost 1.33 billion transistors with up to 16MB of Level 3 cache, Hyperthreading and an unusually high clock speed for a top-end server processor. The FSB800 16MB L3 Xeon MP 7140M part we’re seeing here clocks at 3.4GHz, which is pretty high even for a single core desktop Pentium 4. There also was an FSB667 part called Xeon MP 7150N clocking at 3.5GHz. Only that here we have 2 cores with HT and a metric ton of cache!

These things can run on quad sockets. Meaning a total of 8 cores and 16 threads, like seen on some models of the HP ProLiant DL580 G4. Plus, they’re x86_64 chips too, so they can run 64-Bit operating systems.

Tulsa die shot

Best Tulsa die shot I could find. To the right you can see the massive 16MB L3 cache. There is also 2 x 1MB L2.

And the core point: They’re rare. Extremely rare, especially in the maxed-out configuration of four processors. And I want them tested, as real results are scarce and almost nowhere to be found. Also, Thomsen-XE (who took that photograph of a 7140M up there) wants to see them show off! We have been searching for so long, and missed two guys with corresponding machines by such a narrow margin already!

We want the mightiest of all Netbursts and Intels first native dual core processor to finally show its teeth and prove that with enough brute force, it can even kill the Core 2 micro-architecture (as long as you have your own power plant, that is)!

So now, I’m asking you to please tell us in the comments whether you have or have access to such a machine and if you would agree to run the completely free x264 benchmark on that system. Windows would be nice for a reference x264 result, but don’t mind the operating system too much. Linux and most flavors of UNIX will do the job too! Guides for multiple operating systems are readily available at the bottom of the results list in [English] as well as [German].

If anyone can help us out, that’d be awesome! Your result will of course be published under your name, and there will be a big thank you here for you!

And don’t forget to say bye bye to Tulisa:

Tulisa Contostavlos #1

Well, thanks for your visit, Miss Contostavlos, but TULSA is the #1 we seek today!

Update: According to a [comment] by Sjaak Trekhaak my statements that Tulsa was Intels first native dual core were false. There were others with release dates before Tulsa, like the first Core Duo or the smaller Netburst-based Xeons with Paxville DP core, as you can also see in my reply to Sjaaks comment. Thus, the strike-through parts in the above text.

Jun 132013
 

Wine LogoSure there are ways to compile the components of my x264 benchmark on [almost any platform]. But you never get the “reference” version of it. The one originally published for Microsoft Windows and the one really usable for direct comparisons. A while back I tried to run that Windows version on Linux using [Wine], but it wouldn’t work because it needs a shell. It never occurred to me that I could maybe just copy over a real cmd.exe from an actual Windows. A colleague looked it up in the Wine AppDB, and it seems the cmd.exe only has [bronze support status] as of Wine version 1.3.35, suggesting some major problems with the shell.

Nevertheless, I just tried using my Wine 1.4.1 on CentOS 6.3 Linux, and it seems support has improved drastically. All cmd.exe shell builtins seem to work nicely. It was just a few tools that didn’t like Wines userspace Windows API, especially timethis.exe, which also had problems talking to ReactOS. I guess it wants something from the Windows NT kernel API that Wine cannot provide in its userspace reimplementation.

But: You can make cmd.exe just run one subcommand and then terminate using the following syntax:

cmd.exe /c <command to run including switches>

Just prepend the Unix time command plus the wine invocation and you’ll get a single Windows command (or batch script) run within cmd.exe on Wine, and get the runtime out of it at the end. Somewhat like this:

time wine cmd.exe /c <command to run including switches>

Easy enough, right? So does this work with the Win32 version of x264? Look for yourself:

So as you can see it does work. It runs, it detects all instruction set extensions (SSE…) just as if it was 100% native, and as you can see from the htop and Linux system monitor screens, it utilizes all four CPU cores or all eight threads / logical CPUs to be more precise. By now this runs at around 3fps+ on a Core i7 950, so I assume it’s slower than on native Windows.

Actually, the benchmark publication itself currently knows several flags for making results “not reference / not comparable”. One is the flag for custom x264 versions / compilations, one is for virtualized systems and one for systems below minium specifications. The Wine on Linux setup wouldn’t fit into any of those. Definitely not a custom version, running on a machine that satisfies my minimum system specs, leaving the VM stuff to debate. Wine is per definition a runtime environment, not an emulator, not a VM hypervisor or paravirtualizer. It just reimplements the Win32/64 API, mapping certain function calls to real Linux libraries or (where the user configures it as such) to real Microsoft or 3rd party DLLs copied over. That’s not emulation. But it’s not quite the same as running on native Windows either.

I haven’t fully decided yet, but I think I will mark those results as “green” in the [results list], extending the meaning of that flag from virtual machines to virtual machines AND Wine, otherwise it doesn’t quite seem right.

gt;

May 292013
 

Gainward logoSo, there is this mainboard, an Intel D850EMV2, or rather D850EMVR, which is a sub-version of the former, i850E Rambus chipset. What’s special about that old Pentium 4 board? Well, I won it once in a giveaway at one of the largest german hardware websites, [Computerbase]German flag. And after that, Jan-Frederik Timm, founder and boss of the place contacted me on ICQ, telling me about it. First time I had ever won anything! He asked me to put it to good use, because he was kind of fed up with people just reselling their won stuff. So i promised him that I would.

And boy, did i keep that promise! At first i used a shabby 1.6GHz Northwood processor with just 128MB Rambus RDRAM. Can’t remember the rest of the machine, but over time I upgraded it a bit (when it was already considered old, with Core 2 Duos on the market), a bit more RAM etc. for small LAN party sessions. At some time, I sold it to my cousin, and to make it more powerful and capable for her, I gave the machine the only one and fastest hyper-threaded processor available on that platform, the Pentium 4 HT 3.06GHz, plus 1.5GB PC800-45 RDRAM and my old GeForce 6800 Ultra AGP.

She used that for the internet and gaming etc. for some time until the GeForce died and she thought the machine barely powerful enough for her top-tier games anyway, so she bought a new one, which I built for her. The Intel board and its stuff got my older GeForce FX5950 Ultra then and was used by my uncle for the Internet on Debian Linux and low-end LAN games on Windows XP.

A long time after I first got it, I contacted Jan from Computerbase again, to tell him that I had kept my promise and ensured the board had been used properly for 8 years now. Needless to say he was delighted and very happy that it wasn’t just sold off for quick cash.

Soon after, my cousin got another even more powerful machine, as her Core 2 Duo mainboard died off. Now it was S1156, GTX480 etc. So my uncle bought a new mainboard and I rebuilt the C2D for him with my cousins old GTX275. I asked him if he would part with the D850EMVR and he agreed to give it back to me, after which it collected dust for a year or so.

Now, we need another machine for our small LAN parties, as our Notebooks can’t drive the likes of Torchlight II or Alien Swarm. It was clear, what I had to do: Keep the damn Intel board running until it fucking dies!

This time I chose to make it as powerful as it could remotely become. With a Gainward Bliss GeForce 7800GS+ AGP. The most powerful nVidia based AGP card ever built, equipped with a very overclockable 7900GT GPU with a full 24 pixel pipelines and 8 vertex shaders as well as 512MB Samsung RAM. Only Gainward built it that way (a small 7900 GTX you could say), as nVidia did not officially allow such powerful AGP cards. So this was a limited edition too. I always wanted to have one of those, but could never afford them. Now was the time:

As expected (there were later, more powerful AGP8x systems in comparison to this AGP4x system, with faster Pentium4s and Athlon64s), the CPU is limiting the card. But at least I can add some FSAA or even HDRR at little cost in some games, and damn, that card overclocks better than shown on some of the original reviews! The core got from 450MHz to 600MHz so far, dangerously close to the top-end 7900 GTX PCIe of the time with its 650MHz. Also, the memory accepted some pushing from 1.25GHZ DDR3 to 1.4GHz DDR3 data rate. Nice one!

This was Furmark stable, and the card is very silent and rather cool even under such extreme loads. Maybe it’ll accept even more speed, and all that at a low 1.2V GPU voltage. Cool stuff. Here, a little AquaMark 3 for you:

7800gs+ in AquaMark 3

So, this is at 600MHz core and 1400MHz DDR memory. For comparison I got a result slightly above 53k at just 300MHz core. So as you can see, at least the K.R.A.S.S. engine in AquaMark 3 is heavily CPU bound on this system. So yeah, for my native resolution of 1280×1024 on that box, the card is too powerful for the CPU in most cases. The tide can turn though (in Alien Swarm for instance) when turning on some compute-heavy 128-bit floating point rendering with HDR or very complex shaders, or FSAA etc., so the extra power is going to be used. ;) And soon, 2GB PC1066-32p RDRAM will arrive to replace the 1GB PC800-45 Rambus I have currently, to completely max it out!

So I am keeping my promise. Still. After about 10 years now. Soon there will be another small LAN party, and I’m going to use it there. And I will continue to do so until it goes up in flames! :)

Update: The user [Tweakstone] has mentioned on [Voodooalert]German flag, that XFX once built a GeForce 7950GT for AGP, which was more powerful than the Gainward. So I checked it out, and he seems to be right! The XFX 7950GT was missing the big silent cooler, but provided an architecturally similar G71 GPU at higher clock rates! While the Gainward 7800GS+ offered 450MHz on the core and 1250MHz DDR data rate on the memory, the XFX would give you 550MHz core and 1300MHz DDR date rate at a similar amount of 512MB DDR3 memory. That’s a surprise to me, I wasn’t aware of the XFX. But since my Gainward overclocks so well (it’s the same actual chip after all) and is far more silent and cool, I guess my choice wasn’t wrong after all. ;)

Update 2: Since there was a slight glitch in the geometry setup unit of my card, I have now replaced it with a Sapphire Radeon HD3850 AGP, which gives more performance, slightly better FSAA and as the icing on the cake proper DXVA1 video acceleration. Even plays BluRays in MPC-HC now. ;) Also, I retested AquaMark 3, which seems to require the deletion of the file direcpll.dll from the AquaMark 3 installation directory to not run into an access violation exception at the end of the benchmark on certain ATi or AMD graphics hardware. I guess the drivers are the problem here. But with that troublesome file gone, here’s a new result:

AquaMark 3 on an ATi Radeon HD3850 AGP

Yeah, it’s a bit faster now, but not much. As we can see, the processor is clearly the limiting factor here. But at least I now have relatively problem-free 3D rendering and DXVA on top of it!

Oct 252012
 

Sun Grid Engine LogoOk ok, I guess whoever is reading this (probably nobody anyway) will most likely already be tired of all this x264 stuff. But this one I need as documentation for myself anyway, because the experiment might be repeated later. So, the [chair for simulation and modelling of metallurgic processes] here at my university has allowed me to try and play with a distributed grid-style Linux cluster built by Supermicro. It’s basically one full rack cabinet with one Pentium 4 3.2GHz processor and 1GB RAM per node, with Hyper-Threading being disabled because it slowed down the simulation jobs that were originally being run on the cluster. Operating system for the nodes was OpenSuSE 10.3. Also, the head node was very similar to the compute nodes, which made it easy to compile libav and x264 on the head node and let the compute nodes just use those binaries.

The software installed for using it is the so called Sun GRID engine. I  have once already set up my own distributed cluster based on an OpenPBS style system called Torque, together with the Maui scheduler. When I was introduced to this Sun GRID engine, most of the stuff seemed awfully familiar, even the job submission tools and scripting system were quite the same actually. So this system uses tools like qsub, qdel, qstat plus some additional ones not found in the open source Torque system, like sns.

Now since x264 is not cluster-aware and not MPI capable, how DO we actually distribute the work across several physical machines? Lacking any more advanced approaches, I chose a very crude way to do it. Basically, I just cut the input video into n slices, where n is the number of cluster nodes. Since the cluster nodes all have access to the same storage backend via NFS, there was no need to send the files to the nodes, as access to the users home directory was a given.

Now, to make the job more easy, all slices were numbered serially, and I wrote a qsub job array script, where the array id would be used to specify the input file. So node[2] would get file[2], node[15] would get file[15] to encode etc. The job array script would then invoke the actual worker script. This is what the sliced input files look like before starting the computation:

Sliced input file

And here, the qsub job array script that I sent to the cluster, called benchmark-qsub.sh, the directory /SAS/home/autumnf is the users home directory:

#$ -N x264benchmark
#$ -t 1-19
 
export PATH=$PATH:/SAS/home/autumnf/usr/bin
echo $PATH
 
cd /SAS/home/autumnf/x264benchmark
 
time transcode.sh

And the actual worker script, transcode.sh:

#!/bin/bash
# Pass 1:
x264 --preset veryslow --tune film --b-adapt 2 --b-pyramid normal -r 3 -f -2:0 --bitrate 10000 --aq-mode 1 -p 1 --slow-firstpass --stats benchmark_slice$SGE_TASK_ID.stats -t 2 --no-fast-pskip --cqm flat slice$SGE_TASK_ID.264 -o benchmark_1stpass_slice$SGE_TASK_ID.264
 
# Pass 2:
x264 --preset veryslow --tune film --b-adapt 2 --b-pyramid normal -r 3 -f -2:0 --bitrate 10000 --aq-mode 1 -p 2 --stats benchmark_slice$SGE_TASK_ID.stats -t 2 --no-fast-pskip --cqm flat slice$SGE_TASK_ID.264 -o benchmark_2ndpass_slice$SGE_TASK_ID.264

As you can see, the worker is using the environment variable $SGE_TASK_ID as a part of the input and output file names. This variable contains the job array id passed down from the job submission system of the Sun GRID engine. The actual job submission script contains the line #$ -t 1-19 which tells the system, that the job array consists of 19 jobs, as the cluster had 19 working nodes left, the rest was already dead as the cluster was pretty much out of service and hence unmaintained. Let’s see how the Sun tool sns reports the current status of the grid cluster:

Empty Grid Cluster

So, some nodes are in “au” or “E” status. While I do not know the exact meaning of the status abbreviations, that basically means that those nodes are non-functional. Taking the broken nodes into account we have 19 working ones left. Now every node invokes its own x264 job and gives to it the proper input file from slice1.264 to slice19.264, writing correspondingly named outputs for both passes. Now let’s send the script to the Sun GRID engine using qsub ./benchmark-qsub.sh and check what sns has to say about this afterwards:

Sns reporting a grid cluster under load

Hurray! Now if you’re more used to OpenPBS style tools, we can also use qstat to report the current job status on the cluster:

Qstat showing a cluster under load

As you can see,  qstat also reports a “ja-task-ID”, which is essentially our job array id or in other words $SGE_TASK_ID. So thats basically one job with one job id, but 19 “daughter” processes, each having its own array id. Using tools like qdel or qalter you can either modify the entire job, or only subprocesses on specific nodes. Pretty handy. Now the Pentium 4 processor might suck ass, but 19 of them are still pretty damn powerful when combined, at the moment of writing you can find the cluster on [place #4 on the results list]! Here the Voodooalert style result, just under one hour:

0:58:01.600 | SMMP | 19/1/1 | Intel Pentium 4 (no HT) 3.20GHz | 1GB DDR-I/266 (per node) | SuperMicro/SGE GRID Cluster | OpenSuSE 10.3 Linux (Custom GCC Build)

To ensure that this is actually working, I have recombined the output slices of pass 2, and tried to play that file. To my surprise it worked and would also allow seeking. Pretty nice considering that there is quite some bogus data in the file, like multiple H.264/AVC headers or cut up frames. I originally tried to split the input file into slices at keyframes in a clean fashion using ffmpeg, but that just wouldn’t work for that type of input, so I had to use dd, resulting in some frames being cut up (and hence dropped), and slices 2-19 having no headers. That required very specific versions of libav and x264, as not all versions can accept garbled files like this.

Also, the output files have been recombined using dd. Luckily, mplayer using libav/ffmpeg would play that stuff nicely, but there’s simply no guarantee that every player and/or decoder would. So that’s why it cannot be considered a clean solution. Also, since motion estimation is less efficient for this setup at the cutting points, it’s not directly comparable to a non-clustered run. So there are some drawbacks. But if you would cluster x264 for productive work, you’d still do it kind of like that. Here, the final output, already containing the final concatenated file, quite a mess of files right there:

Clusterrun done

So this is it. The clustered x264. I hope to be able to test this approach on another cluster at the Metallurgy chair in the next months, a Nehalem-based machine with far more cores, so that’d be really massive. Also, access to a Sandy Bridge-E cluster is possible, although not really probable. But we’ll see. If you’re interested in using x264 in a similar approach, you might want to check out the software versions that I used, these should be able to cope with rudely cut up slices quite well:

Also, if you require some guidance building that source code on Linux, please check out my guide:

If anybody knows a better way to slice up H.264/AVC elementary video streams, by all means, let me know! I would love to be able to have slices cut at proper keyframe positions including their own header, and I would also like to be able to reconcatenate slices to one file that is clean, having only one header at the beginning of the file and no damaged / to be dropped frames at their joints. So if you know how to do that – preferrably using Linux command line tools – just tell me, I’d be happy to learn that!

Edit: Thanks to [LoRd_MuldeR] from the [Doom9 Forums] I now have a way of splitting the input stream cleanly at GOP (group of pictures) boundaries, as the Elephants Dream movie is luckily using closed GOPs. Basically, it involves the widely-used [MKVtoolnix]. With that tool, you can just take the stream, split it to n MKV slices, and then either use those or extract the H.264/AVC streams from those slices, maybe using [tsMuxer]. Just make as many as your cluster has compute nodes, and you’re done!

By the way, both MKVtoolnix and tsMuxer are available for Windows and Linux, also MacOS X.

This is clean, safe and proper, other than my dirty previous approach!

Oct 232012
 

DragonFly BSD LogoAt first I thought there were only three big BSD distributions, namely FreebBSD, OpenBSD and NetBSD. But there is a fourth one, which might very well become the last for me to port the x264 benchmark to. DragonFly BSD. Actually a pretty cool system with its own, quite modern file system called “HAMMER”. Surprisingly, like NetBSD it made my life [quite easy], so now 50% of the major BSD systems are covered successfully. OpenBSD might eventually join the ranks as soon as version 5.2 is being released, supposedly including a more modern compiler and hopefully assembler. But so far, it’s NetBSD and DragonFly BSD.

Interestingly, DragonFly BSD not only features its own file system, but also its own light-weight threading implementation. Of course I cannot use that, as I have to stick to a threading model supported by libav/x264 (either BeOS threads which are broken anyway, win32 or posix, so posix it is), but it really looks like a lot of work went into this project. On the down side, it only supports x86_32 and x86_64, where the other BSDs support a sometimes extremely wide range of microarchitectures, with NetBSD clearly [in the lead] even over any Linux distribution.

Now it seems almost everything that can be conquered, has been conquered. From oddballs like the Win32 ReactOS to strange alien systems like Haiku and Unices like Solaris and different BSDs. And Linux of course, on quite some different architectures (MIPS, PPC, ARM, IA-64, …). While OpenBSD might still follow, and someday maybe even FreeBSD, I think the exploration time is pretty much over.

I did try to get on to [Fafner], a VAX running in the cellar of a crazy university professor, but it seems  the OpenVMS operating system with its compilers and build toolchains is far, far beyond my porting skills. Same goes for weirdos like Tanenbaums Minix. So for now, this is it:

DragonFly BSD running x264

Oct 062012
 

ReactOS LogoIt seems that some miracle has happened. I have now tried quite some builds of ReactOS when i stumbled over a certain [v0.4-SVN r57481] just two days ago. Now previously, that OS would just bluescreen on larger data transfers. Older builds would give an NTOSKRNL.EXE bluescreen, newer ones would show the tcpip.sys kernel driver faulting. Also, on x264 CPU/RAM load the kernel would just die after very few minutes, sometimes seconds. So my little x264 benchmark would reach something like 300-400 (of 15691) frames, and then everything would go right down to hell.

As you can imagine, I was quite pleasantly surprised when all of a sudden I could easily download an almost 700MB large file from the internet with v0.4-SVN r57481 without the transfer stalling (which also happened sometimes previously) or the machine bluescreening. I was even more suprised when i ran x264 and found the box to gnaw its way through 8 hours of high load without a single glitch! Now dear developers, that’s quite some progress right there in kernel space, I am impressed!

As for  the x264 benchmark thing, there are a few minor modifications necessary for it to do its time measurement correctly, but that’s no big issue. Links to corresponding guides have already been [added]. Here you can see the thing running for quite some time already, with no crash. I still can’t believe it, so many problems gone all of a sudden:

ReactOS stabilized

Maybe ReactOS can become a real open source replacement for Windows one day? The sudden leap in stability is kind of reassuring. If only they could get more developers working on this stuff.

Sep 252012
 

Haiku LogoAnd so the quest continues, today, a new operating system type: BeOS, in its modern open source form: Haiku. BeOS was originally developed in the 90s as a multimedia operating system to compete with Microsoft Windows and Apple MacOS. However, the operating system never quite took off. In more recent times, the entire OS has been resurrected under a new, japanese name: Haiku.

Still featuring a BeOS style kernel, the OS has been equipped with a lot of Posix & GNU tools to easy the pain of porting and developing software for the OS. It actually looks and feels very tidy, fast and cool.

In recent development builds, there are even versions with GCC 4.6 shipped, but on top of that, what you get is also a complete modern GNU build toolchain, including but not limited to yasm, autoconf, make, imake etc. So I tried to build and link libav and x264 obviously, but failed. One problem was that some OS functions (like Posix thread implementation) are implemented in different libraries (libroot instead of libm or libpthread), so modifications of linker flags are necessary, e.g. -lroot instead of -lm.

But there were more severe problems prohibiting me from linking x264 against libav or ffmpeg itself. I have as of yet not been able to fully figure out why, but it’s most definitely a linker problem with failing library/header detections and even missing references on linking. Maybe some of the libs are even actually missing, but I am not sure why they wouldn’t have built when compiling libav and/or ffmpeg. Well, maybe I’ll manage in the future, meanwhile, check out this Haiku screenshot, it does look rather cool:

Haiku Screenshot

Sep 202012
 

x264 LogoI have played around with PHP a little again, and actually managed to generate PNG images with some elements rendered to them using a few basic GD functions of the scripting language. This is all still very new to me, so don’t be harsh! ;)

I thought I might use this to create some dynamic and more fancy than plain text statistics about the [x264 benchmark]. I decided to do some simple stats about operating systems and CPUs first, didn’t want to overdo it.

So I went for basic OS families and a more broken down visualization of all Windows and all UNIX derivatives. For microprocessors I went for basic architecture families (x86, RISC, VLIW) and a manufacturer breakdown. I know “x86” should probably have been “CISC” instead, but since x86 in itself is so wide-spread, I thought I should just make it its own family. See the following links:

Just so you can see how the generated images look like, I’ll link them in here. As you can see I decided to keep it very plain and simple, no fancy graphics, operating systems first:

Operating systems

Windows operating systems

Windows operating systems

And the microprocessors:

Microprocessor architectures

Microprocessor manufacturers

Not too bad for my first PHP-generated dynamic images? I would sure like to think so. ;)