A high-level, general-purpose programming language, created as an extension of the C programming language, that has object-oriented, generic, and functional features in addition to facilities for low-level memory manipulation.
Processor cache contents cannot be controlled directly from Visual C++ or ISO/IEC 14882:1998 C++ code. The practical approach is to write time-critical code so it is more likely to benefit from cache locality and to avoid patterns that increase cache misses and page faults.
Useful guidance:
- Keep related data together. Data structures with good locality of reference reduce missed cache hits and page faults.
- Prefer arrays over dynamically allocated linked lists when practical. Traversing linked lists can miss the cache or cause page faults because each link may be in a different memory location. In some cases, a simple array-based implementation is faster.
- Be careful with hash tables that use dynamically allocated linked lists. These can perform substantially worse because of poor locality. An array-based hash table (closed hashing) can have superior performance.
- Minimize dynamic allocation overhead in time-critical paths.
For MFC code,
CStringuses dynamic allocation. A simplechararray on the stack can be faster for some scenarios, and constant strings should useconst char *. - If using
CArray, size it up front. UseCArray::SetSizeand specify growth behavior so repeated insertions do not cause frequent reallocations and copies, which can fragment memory and increase cache misses and page faults. - Avoid unnecessary memory overhead in list structures.
CListis a doubly linked list. If a doubly linked list is not required, a singly linked list reduces pointer-update overhead and memory use, which also reduces opportunities for cache misses and page faults. - In parallel code, avoid false sharing.
False sharing happens when separate tasks write to variables on the same cache line, causing repeated cache invalidation and reloads. One mitigation is to place frequently written variables on separate cache lines. When sharing data among tasks,
concurrency::combinableis recommended because it creates thread-local variables in a way that makes false sharing less likely. - Measure before and after changes.
Use Performance Monitor (
perfmon.exe) to gather information about performance.
Also, cache correctness is different from cache performance. On multiprocessor systems, if shared values are accessed concurrently, use proper synchronization. The documented example shows using volatile in /volatile:ms mode or InterlockedExchange to ensure visibility and ordering between processors.