Back in 2000, Quexal changed the way programmers had to deal with MMX programming. A friendly user interface simplified building parallel versions of algorithms, an optimizing compiler made sure that the resulting code would run fast, and a visual debugger helped pinpoint programming errors. Even if the focus of the…
-
-
The Contact form in Joomla! uses the mail server settings to send the contact request email to the site administrator. However, when using SMTP servers (I’ve tried both Yahoo! and GMail SMTP servers), it does not work, as the email address of the sender is not recognized by the SMTP…
-
AltaLux is an image processing technology that can significantly enhance the quality of images and videos with poor lighting conditions. AltaLux/Demo is a Windows sample application that lets you enhance the quality of JPEG images for free (download it! from the download area). For even better usability, I strongly recommend…
-
A new article about using Intel TBB is here. It contains examples using C++ lambdas and joining multi-threaded loops with SIMD code In this article we will transform a plain C loop into a multi-threaded version using Intel Thread Building Blocks library (TBB). Here is the loop to transform: {CODE…
-
Profile Software engineering leader with more than 20 years of experience, including over a decade managing teams. Led distributed teams of up to 14 developers at AWS RDS and Amazon Robotics, delivering cloud and robotics platforms with measurable gains in cost, reliability, and operational efficiency. Pairs people leadership and product…
-
SSE2 extends the original SSE instruction set with support for double-precision floating-point arithmetic and a wider set of integer SIMD operations. The original SSE instructions operate mainly on four 32-bit single-precision floating-point values stored in a 128-bit XMM register. SSE2 adds the ability to operate on two 64-bit double-precision floating-point…
-
SSE, short for Streaming SIMD Extensions, is an x86 instruction set extension that allows one instruction to operate on multiple floating-point values at the same time. The most common SSE programming model uses 128-bit XMM registers. Each register can hold four 32-bit single-precision floating-point values: Instead of adding one float…
-
MMX was Intel’s first widely used SIMD extension for x86 processors. It introduced packed integer operations, allowing one instruction to process multiple small integer values at the same time. MMX is mostly historical today, but it is still useful to understand older multimedia, image-processing, audio, codec, and game code. Many…
-
SSE arithmetic lets one instruction operate on four 32-bit floating-point values packed in a 128-bit XMM register. This page explains the arithmetic instructions, their C/C++ intrinsics, the difference between packed and scalar forms, and the practical pitfalls that matter when writing or reading SSE code. Packed vs scalar: PS and…
-
MMX was Intel’s first widely adopted SIMD instruction set on x86 processors. It introduced 64-bit packed integer operations and made it possible to process multiple bytes, words, or doublewords with a single instruction. For the late 1990s, MMX was an important step forward. It was useful for image processing, audio…