When conventional filesystem or structural analysis does not provide a clear starting point, searching raw data for recognizable patterns or file signatures can help identify potentially relevant artifacts. Data carving can be applied to a variety of raw data sources, including unallocated and slack space, swap or pagefile data, and memory images.
Carving memory images for whole files, however, is generally low-yield. Most file carvers rely on headers and either footers or a maximum size (or, for structure-aware carvers, format-specific length fields and validators). Footers are especially unreliable in memory. The core problem is the mismatch between file size and page granularity. On x86/x64 and most ARM configurations, the base page size is 4 KB, though this is not universal: Apple operating systems on Apple Silicon commonly use 16 KB pages, some ARM64 Linux configurations support 16 KB or 64 KB pages, POWER commonly uses 64 KB, and Windows and Linux can back some allocations with 2 MB large pages. Some artifacts of interest are smaller than a page (LNK files, small scripts, configuration data), but most files that matter in an investigation (Office documents, PE files, images, archives) span multiple pages.
Contiguity in a process's virtual address space does not imply contiguity in physical memory. A virtually contiguous region may be backed by physically non-contiguous pages. Some virtual pages may also be absent from a physical-memory acquisition because they were paged out, never populated, or otherwise not resident at acquisition time. Consequently, a header-to-footer carve across page boundaries in a raw physical-memory image can produce truncated, incomplete, or incorrectly assembled data. Unlike a filesystem, physical memory does not provide a file-allocation mechanism that preserves the physical contiguity of a file's pages.
For recovering whole files from memory, structured recovery is therefore often preferable to traditional signature carving. In Windows memory images, for example, Volatility's filescan can be used to locate FILE_OBJECT structures, while dumpfiles can attempt to recover cached file data using the Windows cache and section-related structures, including the SharedCacheMap. Recovery remains dependent on the relevant structures and pages being present in the captured memory image, so gaps and incomplete files should be expected.
Carving memory does pay off when the data you are looking for fits within a single page or is otherwise self-contained: URLs, IP addresses, command lines, configuration blobs, encryption keys, credentials, and other small records. These survive intact regardless of physical fragmentation.
In this post, we will examine several carving techniques, including page_brute, a tool that applies YARA rules to data page by page (a design that assumes 4 KB pages and is most useful against pagefile and raw page data) and several Volatility Framework plugins that perform unstructured data carving, such as strings.
GNU strings
The strings utility extracts printable character sequences from binary data. It is built into Linux/Unix (GNU binutils); on Windows, the Sysinternals Strings utility provides similar functionality with different options. The options discussed below are those of GNU strings.
You can run strings against memory images, extracted memory sections, pagefiles, or PE files to identify potentially useful artifacts and infer aspects of a program's functionality. Common targets include IP addresses, URLs, registry keys, mutex (mutant) names, command lines, filenames, and imported or referenced API names. Together with grep, it is an essential examiner skill. Its output can broaden the scope of an investigation, for example, when a known malware mutex name appears in a memory image or pagefile. However, the presence of a string alone does not establish that the associated code executed: strings can originate from unrelated sources such as security tooling, browser caches, documentation, embedded resources, or another unrelated process. Strings should therefore be corroborated by attributing them to a process or memory region and by examining other evidence.
By default, GNU strings reports sequences of four or more printable 7-bit characters (-e s). Windows memory images contain both single-byte ASCII and UTF-16 strings (Windows uses little-endian UTF-16 internally). Consequently, when examining Windows memory, it is useful to perform at least two passes: one default pass for single-byte strings and one using -e l for 16-bit little-endian characters. The -e b option searches for 16-bit big-endian sequences, while -e S, -e L, and -e B support 8-bit and 32-bit character encodings. These modes perform pattern-oriented extraction rather than full Unicode-aware text decoding, so non-Latin or otherwise complex UTF-16 text may not be displayed as expected.
The minimum string length can be changed with -n. Increasing the threshold can substantially reduce noise when scanning large memory images. The -a option instructs GNU strings to scan the entire file rather than restricting the search to the initialized and loaded sections of an object file. The -t option records the offset at which each string was found: use -t x for hexadecimal, -t d for decimal, or -t o for octal.
When strings is run against a straightforward raw physical-memory image, these offsets identify locations within the memory-image file and generally correspond to physical-memory offsets; they are not process virtual addresses. To determine whether a string belongs to a particular process or virtual address space, correlate the location with the operating-system memory structures using a forensic framework such as Volatility. A physical page may also be mapped into more than one address space, so a physical offset should not automatically be treated as belonging to a single process. A typical workflow is:
strings -a -t d mem.raw > ascii.txt
followed by:
strings -a -e l -t d mem.raw > unicode.txtThe resulting files can be merged or sorted by offset and searched with tools such as grep. For example, an examiner might search for IP addresses, registry paths, suspicious filenames, command-line fragments, or known indicators associated with a malware family.
Finally, strings only reveals character sequences that are sufficiently represented in the scanned data. Packed binaries, dynamically resolved APIs, stack-built strings, and encoded or encrypted strings may not be recovered by a straightforward scan. Consequently, sparse output or an unexpectedly sparse set of recognizable strings can itself provide a useful investigative clue, but it should not be treated as proof that a binary or memory region lacks functionality.
GREP Pattern Matching
After generating strings output, the next step in unstructured analysis is often to search it for patterns of interest with grep (from “global regular expression print,” associated with the g/re/p command in ed). GNU grep is commonly available on Linux/Unix systems, including the SIFT Workstation. Windows does not ship with GNU grep, although findstr and PowerShell's Select-String provides a more limited built-in alternative.
grep reads a file or standard input and prints lines that match a specified pattern. This works particularly well with strings output because each extracted string is normally represented as a separate line. Searching raw binary data is different: binary data does not have meaningful text-line boundaries, and grep may report a "Binary file matches " message rather than displaying matching content. The -a option forces GNU grep to treat the input as text, which can be useful for simple searches of raw binary data, but it does not make the binary data equivalent to properly extracted strings. Because Windows memory commonly contains both ASCII and UTF-16LE strings, search the output of both strings passes when investigating Windows memory. Otherwise, important artifacts—such as mutex names, registry paths, filenames, or text from applications such as Notepad—may be missed because they are represented as UTF-16 little-endian character sequences.
We will use grep when scanning process memory dumps for mutex (mutant) strings, IOCs, and text recovered from processes like Notepad. Structural anomalies such as hooked SSDT or IDT entries are found with dedicated Volatility plugins, not by grepping strings.
Useful options include:
-i: perform case-insensitive matching.-v: invert the match and display lines that do not contain the pattern.-c: count matching lines rather than individual occurrences. To count matched occurrences,grep -o pattern file | wc -lcan be used.-n: display line numbers.-A,-B, and-C: display lines after, before, or around a matching line.-o: print only the portion of each line that matches the pattern.-w: match whole words.-E: use extended regular expressions.-P: use Perl-compatible regular expressions where supported by the installed GNUgrep.-F: interpret the search pattern as a fixed string rather than a regular expression. This is useful for exact IOCs, such as IP addresses, because characters such as.otherwise have special meaning in regular expressions.-f file: read multiple search patterns from a file, one pattern per line, which is useful for IOC sweeps.
By default, grep uses basic regular expressions and performs case-sensitive matching. Setting LC_ALL=C makes matching use the C/POSIX locale, avoiding locale-dependent character handling on non-UTF-8 data and often improving performance during searches on large images. It does not, however, make arbitrary binary data safe or meaningful to process as text. To preserve image offsets in results, generate the strings output with an offset option, such as -t d or -t x before grepping. Be aware that grep -b reports byte offsets within the strings output file; those offsets are not offsets into the original memory image. The original memory-image offsets should therefore be preserved by strings itself.
For example:
strings -a -t d mem.raw > ascii.txt strings -a -e l -t d mem.raw > unicode.txt
The resulting files can then be searched with commands such as:
grep -i "suspicious_mutex" ascii.txt
grep -i "suspicious_mutex" unicode.txt
or, for an exact IOC:
grep -F "192.0.2.10" ascii.txt
A match should be treated as an investigative lead rather than proof of execution or attribution. Where possible, correlate the matched string with its memory location and the process or address space in which it occurs using a memory-forensics framework such as Volatility, and corroborate the finding with other evidence.
grep can print lines surrounding each match. The -A n option prints n lines after the match, -B n prints n lines before it, and -C n prints n lines before and after. These options are especially useful for Volatility plugins that dump multi-line records of kernel object fields. For example, to find threads that were on-CPU when the memory was acquired, search the threads output for State: Running. In each record, the ETHREAD: ... Pid: ... Tid: ... line precedes the state line, so use -B with a count large enough to reach it (about 6 lines, depending on plugin version), for example vol.py -f mem.raw --profile=<profile> threads | grep -B6 "State: Running". The exact number of preceding lines depends on the plugin version and output format. -C also works, but it includes trailing lines that may not be necessary when the objective is simply to identify the thread associated with the matching state.
Interpret the result carefully. Memory acquisition takes time, so the captured state of running threads is subject to acquisition-time inconsistency, sometimes referred to as memory-image smear. A logical processor can execute only one thread at a time, so at most one thread per logical CPU can be in the running state at a particular instant. Threads belonging to the acquisition mechanism may be among the running threads, but other system or application threads may also be running when the relevant memory pages are captured. When multiple non-contiguous matches are found, GNU grep separates their context groups with --; overlapping context regions are merged rather than printed repeatedly. These behaviors are useful to remember when interpreting the output as records rather than as independent matching lines.
Volatility 3 plugins generally produce structured or tabular output, making fixed context options such as -A, -B, and -C less useful for record reconstruction. When a text-producing plugin emits records of variable length, parsing the record separator can be more reliable than assuming a fixed number of lines. For example, if the plugin uses a sequence of hyphens as its record separator, an awk expression such as:
awk -v RS='------' '/State: Running/'
can return the complete records containing the desired field. The exact record separator should be adjusted to match the output produced by the specific plugin and version being examined.
This example uses -B1 for grep. We want to capture the command line (by matching "Command line"), and also the preceding line containing the process name and PID, using -B1. This works well against Volatility 2's cmdline or dlllist plugin output, where the PID line sits immediately above the command line with no line between them. The exact layout is plugin- and version-dependent, however. If the output contains a blank line, banner, or other intervening text, a larger context value may be necessary. For variable-length records, parsing the plugin's record separator can be more reliable than relying on a fixed line count.
A second use case: -C5 used to capture PE version resource fields (VERSIONINFO/RT_VERSION). Here, we match on "InternalName" and use -C5 to also capture nearby fields such as CompanyName, ProductName, FileDescription, and FileVersion. Since VERSIONINFO is a small, fixed set of fields, increasing the context count beyond that block mostly pulls in unrelated surrounding strings rather than additional version data. Choose a context value appropriate to the output being examined rather than simply making it as large as possible.
EVTXtract is a Python fragment carver for Windows Event Log (EVTX) data, written by Willi Ballenthin. Since Vista, Windows event logs use a Microsoft-specific binary XML format, which organizes log data into 64-KB chunks. Each chunk begins with a well-known signature header, "ElfChnk\x00". EVTXtract uses this signature to identify the start of intact chunks, but its main value is that it can also locate individual event records directly, using each record's own signature, independent of whether a full chunk header is present. This matters because unallocated space, pagefiles, and memory images frequently contain partially overwritten or fragmented chunks, where the chunk header is gone but individual records, or portions of them, survive.
Once the tool applies sanity checks to filter out false-positive matches, it extracts records and recovers templates. Because the binary XML format stores a record's structure separately, as a template referenced by GUID, with the record itself holding only substitution values, EVTXtract must correlate records against templates that may be recovered from elsewhere in the same chunk or even a different chunk or file, rather than assuming each record is self-contained. It decodes this substitution data and reconstructs full event records from it, making it useful for recovering deleted, corrupted, or overwritten event log entries.
evtxtract -v <memory_image>This carved event entry is from the Microsoft-Windows-Bits-Client/Operational log. BITS (Background Intelligent Transfer Service) is a Windows component most often associated with legitimate uses such as Windows Update, but it can also be invoked directly (via bitsadmin or Start-BitsTransfer) by any process, including malware, to move files under the guise of a trusted system service. In this case, the record shows a BITS transfer job (Transfer ID/ID fields) named BITS Transfer @, targeting a URL at dasmalwerk.eu/zippedMalware/...7z — a .7z archive, hosted on a known malware-sample site rather than a Microsoft update endpoint. This pattern is consistent with BITS being abused to download a malicious payload rather than with normal Windows Update activity, and the zero values in Bytes Transferred and Bytes Transferred From Peer indicate the transfer had not yet moved data at the point this record was captured.
Unstructured analysis applies broadly in memory forensics: it works effectively against essentially any raw data, including full memory images (and disk images), the system pagefile, the hibernation file, and extracted process memory sections, since it doesn't rely on knowing or parsing the underlying structure.
One common IR use case is dumping the memory of the process responsible for console/terminal handling to recover attacker command-line activity, including commands from cmd.exe sessions that have since terminated. In Windows XP and Vista, csrss.exe (Client/Server Runtime Subsystem) was responsible for console window management, so dumping its memory and running strings against it can reveal commands entered in cmd.exe sessions. Starting with Windows 7, this responsibility moved to conhost.exe, introduced specifically to move console I/O handling out of the highly privileged csrss.exe process for security reasons. A conhost.exe process is generally tied to a console session and can service more than one client console application at once, though it is still commonly associated with a given cmd.exe window. Because conhost.exe processes are created and torn down with each console session, rather than persisting for the life of the system session the way csrss.exe does, less historical command data tends to remain resident in memory on Windows 7 and later compared to XP/Vista systems.
Another significant target for unstructured analysis is the system pagefile, which frequently holds unstructured data directly relevant to an investigation, as covered in more detail later.
Bulk Extractor
One of the most impressive, efficient, free, and open-source tools for unstructured data analysis is Bulk Extractor, written by Simson Garfinkel, associated with the Naval Postgraduate School in Monterey, CA. Bulk Extractor can be run against any data regardless of filesystem type, or even with no filesystem at all, as is the case with physical memory images. It is multithreaded, so it can make full use of multiple cores. It parses data based on pattern matching, extracting things like network packets, email addresses, credit card numbers, and, via a dedicated scanner, AES encryption keys, into separate feature-report files.
Another distinguishing feature is its ability to decompress data it finds on the fly and recursively re-scan the decompressed stream, so encoded or embedded content that traditional carving tools miss still gets examined. This includes zip, gzip, and rar archives (which also covers the zip-based Office Open XML format used since Office 2007), as well as Windows hibernation files, which use XPRESS compression and are handled by a dedicated hibernation-file scanner.
Bulk Extractor was first publicly released in 2007, with a rewritten, multithreaded version released in May 2010. It has gone through several revisions since, adding functionality such as a portable executable carver and a prefetch file parser.
Bulk Extractor is a multithreaded carving program that accepts a disk image, file, or directory of files as input. Rather than looking for headers and footers, it searches for small, validatable "shards" of structured data. Credit card numbers are a good example: most follow issuer-specific length conventions (commonly but not always 16 digits), the leading digit(s) are constrained to known issuer ranges, and the full number can be checked against the Luhn checksum. Other data types, like a U.S. phone number (three digits, dash, three digits, dash, four digits), are matched via regular expression. The program runs a number of targeted scanners to parse whatever data types are present, and it makes no attempt to use filesystem structures — one of its real strengths, since it keeps working even when a filesystem is damaged, corrupted, partially deleted, or, as with a memory dump, absent entirely.
Bulk Extractor also recovers a wordlist, enabled with -e wordlist (disabled by default). It tokenizes and deduplicates words found in the target data into a dictionary useful for password cracking — a related but distinct output from simply running strings against the data, since programs often fail to clear buffers holding a user's password, so that material can end up in the wordlist. The wordlist can be browsed in BEViewer, the GUI report-viewing tool included with Bulk Extractor, and a version of it can be fed directly to password-cracking tools.
Other scanners not enabled by default include:
-e base16: enables the base16 scanner-e facebook: enables the facebook scanner-e outlook: enables the outlook scanner-e sceadan: enables the sceadan scanner-e xor: enables the xor scanner
Bulk Extractor also includes a Windows PE scanner, scan_winpe, and an ELF scanner, scan_elf, which extracts Executable and Linkable Format files, found largely on UNIX/Linux systems. It can also pull FAT and NTFS directory entries from target data using the windirs scanner (scan_windirs), and it carves for rar, zip, JPEG, GPS coordinates, and Windows LNK files.
One of the most powerful features of Bulk Extractor is its capability to recursively process fragments. The program can recognize fragments that are compressed — for example, parts of a Windows hibernation file or a zip file. The content inside such a fragment cannot be recognized directly by pattern matching while it remains compressed. A compressed phone number, for instance, won't match a phone-number regex, since compression destroys the plaintext byte patterns the regex depends on. A structure-aware carver might still recognize and extract the compressed container itself (the zip member, say), but neither it nor a plain regex scan can recognize the phone number inside that container until it is decompressed — an examiner relying on either approach alone would miss it.
When Bulk Extractor detects compressed data, it decompresses it and recursively re-examines the decompressed result, feeding it back into the pool of data to be scanned by every other scanner. This recursion continues down to a maximum depth controlled by the -M/--max_depth option (default varies by version — check the specific release in use rather than assuming a fixed number). The image below shows an example of this in action for a zip file fragment: the fragment is decompressed, and the resulting data is added back into the chunks to be examined.
Note that although recursive processing works well on disk images, it tends to work less well on memory images, where files are much less likely to be stored contiguously across physical memory, so compressed streams are more likely to be broken up in ways that prevent clean decompression.
Along with recovering discrete features, Bulk Extractor generates a histogram for each feature type it extracts — a frequency count of occurrences, written to files such as email_histogram.txt, with entries formatted as n=<count> followed by the feature value, sorted in descending order of occurrence. Bulk Extractor's documentation notes that frequent features can help an examiner identify patterns involving users, organizations, correspondents, services, and other aspects of system activity.
For email addresses, a highly frequent address may provide a useful investigative lead concerning the identity or communications associated with the examined media. However, frequency alone does not establish that the address belongs to the primary account holder. The same address could appear repeatedly because it belongs to a correspondent, service account, automated process, cached content, or another unrelated source. The same reasoning can be applied to memory images, but with additional caution. A memory capture represents a transient subset of system state rather than the much larger and more persistent body of data typically available from a storage image. Consequently, the relative frequency of candidate addresses in memory may be influenced substantially by what happened to be resident in memory when the acquisition occurred.
In the figure above, benbitdiddle@hotmail.com has the highest frequency, with n=841. The appropriate factual finding is that Bulk Extractor recorded 841 occurrences of that normalized email feature in the processed input. That measurement does not, by itself, establish ownership, custody, or use of the account. The frequency may justify a working hypothesis that the address is significant to the examined system, but that hypothesis should be tested against other evidence. Possible explanations include the address belonging to the subject, a correspondent, a service, or another person or application whose data was present on the media. Corroborating artifacts—such as mail-client configuration, account records, stored credentials, messages, login information, or other attributable evidence—can help distinguish among these possibilities.
Consistent with sound forensic report-writing practice, findings should be documented with precision commensurate with the underlying evidence: state what was recovered and the measurable fact that supports it, without extending the conclusion beyond what the data independently establishes. Interpretive or opinion-based conclusions fall outside the scope of a standard examination report and are reserved for circumstances in which the examiner is serving in an expert-witness capacity or has been specifically tasked with rendering an expert opinion.
Users can also search for custom regular expressions without writing a scanner. The -f <pattern> option accepts a single RE2 regular expression, while -F <file> reads multiple patterns from a file. Matching is case-insensitive by default; --find-case-sensitive can be used when case-sensitive matching is required. Matches are reported by the find scanner to find.txt. Writing a full custom scanner plugin, which involves compiling new C++ extraction logic, is the more involved path, but ad hoc pattern searches are straightforward. Bulk Extractor is primarily developed for Unix-like environments such as Linux and macOS. Windows is supported through a maintained MinGW cross-build and Windows execution test, but native Windows builds are not supported in the current project. The current build process produces a Windows CI artifact rather than a native Windows release installer. The project's release cadence has varied over time, so claims about how frequently it is updated should be checked against the current GitHub repository and release history. Version 2.1.0 was released in January 2024, while the current project documentation describes bulk_extractor 2.2.
The help screen lists many command-line options. In this post, we will use only a few of them. To extract a wordlist, explicitly enable the wordlist scanner with -e:
-e wordlist
You also need to specify an output directory with -o:
-o output
bulk_extractor creates the specified output directory when necessary. With current versions, an existing output directory can also be used to resume an interrupted scan when it contains the appropriate previous-run data. If you want a fresh scan, use a new output directory or follow the version-specific procedure for resetting an existing output directory.
To reduce false positives—such as common email addresses or other frequently occurring features—a stop list can be supplied. Features matching entries in the stop list are diverted from the normal feature output, allowing the examiner to focus on more relevant results. Current documentation notes that stopped results are retained in corresponding _stopped.txt files, so they are not necessarily discarded completely.
-w stoplist.txt
Finally, you can specify the number of worker threads used for analysis with -j:
-j #
bulk_extractor -e wordlist -w stoplist.txt -j 4 -o output image.raw
The exact scanner set, defaults, and available options depend on the installed bulk_extractor version. Use bulk_extractor -H to inspect the scanners and scanner-specific settings available in your build, and bulk_extractor -h for the current command-line options.
We are now going to execute Bulk Extractor against a memory image and review the resulting output together. The output will be written to the analysis/output directory.
As you will observe, Bulk Extractor produces verbose console output and may report warnings or errors while attempting to identify, carve, and decompress artifacts from the memory image. Some errors are expected when examining raw memory and do not necessarily indicate a problem with the tool or the acquisition. Memory images can contain incomplete structures, arbitrary byte sequences that resemble file or compression signatures, and data that cannot be reconstructed completely. Physical memory also does not preserve the logical layout of a process's virtual address space. A process's contiguous virtual address range may be backed by non-contiguous physical pages, while some virtual pages may have no corresponding resident physical page at acquisition time—for example, because they were paged out, had not been populated, or had not been accessed. Consequently, a scanner operating directly on a physical memory image may encounter incomplete or discontinuous byte sequences. Attempts to parse or recursively decompress such data can therefore produce partial results or expected parsing/decompression errors.
Bulk Extractor processes input in pages with an overlapping margin specifically to help scanners recognize features that span page boundaries. Nevertheless, incomplete data, corrupted structures, missing pages, or invalid candidate signatures can still cause individual carving or decompression attempts to fail.
Total runtime is influenced by several factors. Worker-thread allocation (-j) and the amount of data being examined are important factors, but runtime also depends on the scanners enabled, the amount of recursive processing performed, the configured recursive-depth limit (-M), page and margin settings, and the read/write performance of the storage hosting the image and output. The current Bulk Extractor documentation specifically notes that performance depends on the image, storage, scanner set, and recursive workload. Examiners should therefore consider these variables when estimating processing time or investigating an unexpectedly long or short run.
Let us look through the data we found.
The first file we will look at is the AES keys file. You can view this file using the cat command.
This output is useful for demonstrating an important point about the aes scanner in Bulk Extractor 2.0.5: the entries should be treated as candidate AES key material, not automatically as confirmed encryption keys. Bulk Extractor documents the aes scanner as detecting in-memory AES keys from their key schedules. Each record contains three main fields:
70818464 a6 92 06 e5 5d d6 cd d3 86 56 c6 d2 9f 1a 7e a9 AES128
│ │ │
│ │ └─ detected AES key type
│ └──────────────────────────────────────────────────── candidate key material
└──────────────────────────────────────────────────────────────── offset- 70818464 — the offset in the input memory image where the candidate was detected.
- Hexadecimal byte sequence — the candidate AES key material reported by the scanner.
- AES128 / AES256 — the AES key size inferred by the scanner.
The AES scanner does not simply search for arbitrary 16- or 32-byte sequences. It searches for structures consistent with AES expanded key schedules and uses the relationships between words in a candidate schedule to identify likely AES keys. This approach is related to techniques developed for recovering AES keys from memory in cold-boot research.
1. Repeated key values at different offsets.
Several candidate values occur multiple times. For example, the AES-256 candidate beginning 2d 7f 50 78 21 67 9e 90... occurs at several offsets, including 400762224, 583378288, and 884061552. Repeated candidates are worth investigating because they may represent multiple copies of the same cryptographic material in memory. However, repetition alone does not prove that the candidate is a valid operational key or explain why the value occurs at those locations.
2. Clusters of candidates.
Some detections occur in relatively close groups. For example, candidates occur around 400762224–400763544 and again around 583378288–583379608. These clusters are interesting because they may reflect related cryptographic state or repeated memory structures. However, the aes_keys.txt output alone isn't enough to determine the exact data structure each cluster represents.
3. Context matters.
Because the input is a raw physical memory image, an offset in aes_keys.txt does not by itself identify the process that owned the memory. Process attribution requires correlation with the Windows memory-management structures and the relevant process address space. If the objective is to investigate AsyncRAT, candidates should therefore be correlated with the AsyncRAT process and with other artifacts from the same memory image.
Caveats
- AES key-schedule detection is heuristic; a reported candidate is not cryptographic proof that the corresponding bytes were being used as an AES key.
- Multiple occurrences of the same candidate are useful corroborating evidence but do not independently establish validity.
- A candidate becomes substantially more significant when it can be correlated with the process, application, configuration data, or ciphertext for which it was actually used.
- The aes_keys.txt output should therefore be treated as a lead for further analysis, rather than as a definitive list of confirmed encryption keys.
In this example, the presence of numerous AES-256 candidates is interesting in the context of a memory image associated with an AsyncRAT investigation, but the output alone does not establish which candidate, if any, was used by AsyncRAT for configuration or C2 encryption.
One especially important distinction for the lesson is: “Bulk Extractor found an AES key candidate” ≠ “we recovered the AsyncRAT encryption key.” That distinction is worth emphasizing to examiners because it prevents them from treating scanner output as automatically validated forensic evidence. Bulk Extractor itself describes the AES output as the result of detecting in-memory AES keys from their key schedules.
The next file we examine is wordlist.txt. It contains words extracted from the input, somewhat analogous to running strings and then applying length filtering. In the version used for this exercise, the scanner extracts ASCII strings within its configured minimum and maximum word lengths. These limits can be adjusted with the wordlist scanner's configuration options; because scanner-specific settings can vary between Bulk Extractor releases, use bulk_extractor -H to confirm the defaults and available settings for the installed version.
Viewing wordlist.txt opens it through a paging program such as less, allowing you to scroll through the results. The spacebar advances one page, Page Up/Page Down move through the file, and q exits the pager. Consult the command-line cheat sheet for additional less functionality, including searching within the displayed file.
Each feature record contains the offset where the word was found followed by the extracted word. The wordlist can be very large, and unlike the AES scanner which runs by default, the wordlist scanner is disabled by default in Bulk Extractor; it must be explicitly enabled with:
-e wordlist
Bulk Extractor can also produce a wordlist_histogram.txt file, which summarizes the frequency of extracted words. The exact set of output files can vary by Bulk Extractor version and scanner configuration, so confirm the files produced by your installed version rather than assuming that every release generates the same output.
Because wordlist.txt can contain a very large number of entries, reviewing it manually is impractical. Searching it with tools such as grep is much more efficient. For example, memory may contain environment-variable strings associated with processes. USERNAME can be an interesting search term because Windows commonly uses it to represent a user's account name. To search for the literal text USERNAME=:
$ grep "USERNAME=" wordlist.txt
Remember: a USERNAME= hit in a memory image is a lead, not automatically proof of the currently logged-on user. Correlate it with process context and other Windows artifacts before drawing that conclusion.
The grep program is case-sensitive by default. If you want to do case-insensitive searching, use the -i flag. You can see this in action by searching for win10 with and without the -i flag.
Let's turn our attention to the email addresses recovered by Bulk Extractor. The relevant feature file can be viewed with:
$ less email.txt
You will find a substantial number of recovered email addresses in this file. Each feature record contains the offset at which the feature was located, the recovered email address, and surrounding context bytes. The context provides a configurable window of bytes surrounding the detected feature; the default size and available -C settings depend on the Bulk Extractor version and configuration.
The context field is particularly useful when evaluating the significance of a recovered address. For example, an address appearing near recognizable email-related data such as From:, To:, or Subject: may provide stronger contextual support for an email-related interpretation than the same address appearing in isolation or within an unrelated binary structure.
However, context should be treated as supporting evidence rather than proof of provenance, use, or ownership. A memory image can contain stale data, application strings, cached content, copied buffers, and other artifacts that are no longer actively associated with the user or process that originally generated them. The examiner should therefore consider the recovered address together with its surrounding context and other corroborating artifacts before assigning interpretive significance to the finding.
Upon reviewing the recovered email features and their frequency, it becomes apparent that some of the most frequently occurring addresses are unlikely to represent addresses the user actually entered, communicated with, or was otherwise meaningfully exposed to. cdsteam@microsoft.com is a representative example: an address of this type may occur as a built-in or application-generated artifact within Microsoft software. Its repeated presence in memory may therefore reflect the presence or operation of the relevant software component rather than user activity involving that address.
Artifacts of this nature can be suppressed through the use of a stop list. Bulk Extractor's -w option accepts a stop-list file and diverts matching features from the normal feature output. Importantly, current Bulk Extractor versions retain stopped results in the corresponding _stopped.txt files, allowing the examiner to audit what was suppressed. A stop list therefore reduces noise during examination; it does not establish that a suppressed feature is irrelevant or permanently remove it from the results.
A reference stop list of common system- and application-generated artifacts has historically been distributed with Bulk Extractor materials. If using an externally hosted stop list, the examiner should verify that the resource is still available, determine which Bulk Extractor release it was designed for, and review its contents before incorporating it into an examination workflow. Stop lists should be treated as case- and version-dependent analytical aids, not as universal exclusions.
In particular, an examiner should preserve the original unsuppressed results and document the stop-list version and configuration used. This makes it possible to reproduce the analysis and distinguish between an artifact that was absent from the evidence and one that was present but intentionally diverted from the primary feature file.
The same limitation applies to the other types of data we are about to examine. Recovered URLs, including search-related URLs, should not automatically be interpreted as URLs that the user typed or intentionally visited. Browsers and other applications routinely place URLs and related content into memory as pages are loaded and processed. For example, some social-media URLs may have been loaded into memory simply because links to those sites were present on a webpage being viewed. Their presence therefore does not, by itself, demonstrate that the user visited those sites, searched for them, or interacted with the corresponding accounts. Proceed with caution: treat recovered URLs as artifacts requiring contextual interpretation and corroboration, rather than as direct evidence of user intent or activity.
There are several pertinent feature files related to URLs recovered from the examined media. The primary output, url.txt, contains URLs identified by Bulk Extractor's URL scanner. As with the other feature types discussed previously, the volume of entries can be substantial, making sequential manual review impractical. Examiners are therefore better served by beginning with domain- or service-level summaries, such as url_services.txt, which summarizes recovered URLs by service or domain, and domain_histogram.txt, which provides broader domain-frequency information, before examining individual entries in url.txt.
A critical evidentiary caveat applies to all URL-related artifacts recovered from a memory image: presence does not establish access or user intent. A recovered URL may have appeared as part of rendered webpage content, an unclicked hyperlink, a referrer, cached or prefetched content, or data embedded within an application's executable resources. None of these circumstances, by itself, establishes that the subject navigated to the address or intentionally interacted with it. This distinction must be preserved when interpreting and reporting URL-related findings.
Beyond url.txt, the URL scanner in the configuration used for this examination produces several specialized URL-derived feature files. These include url_facebook-address.txt, containing recovered Facebook address/vanity-style URLs; url_facebook-id.txt, containing numeric Facebook identifiers recovered in forms such as id=<number>; url_microsoft-live.txt, containing identifiers associated with Microsoft Live services; url_searches.txt, containing a histogram of probable search terms extracted from URLs associated with search services such as Google, Bing, and Yahoo; and url_services.txt, which provides service/domain-oriented URL information.
Regarding url_facebook-id.txt, numeric Facebook identifiers should be treated primarily as identifiers requiring further correlation, rather than as direct evidence of a particular profile or account. Historically, numeric identifiers could be incorporated into Facebook URLs and resolved by the service, but Facebook's URL and identifier behavior has changed over time. Examiners should therefore verify the behavior applicable to the relevant examination period rather than relying on assumptions derived from older documentation.
Regarding url_searches.txt, recovery of an apparent search term does not establish that the subject actually executed that search. The term may have originated from an unclicked link, a referrer, cached or application-generated content, or other browser data rather than from an affirmative search action by the user. Search-engine URL formats also change over time. Consequently, scanner limitations documented for historical URL formats should not automatically be treated as limitations of the current scanner; when such limitations matter to an examination, they should be verified against the Bulk Extractor version and configuration being used.
Network packets represent data transmitted or received through a network interface. Depending on the protocol stack, this traffic may contain link-layer, IP, TCP, UDP, ICMP, or application-layer structures. During network activity, portions of this data and associated networking structures may reside temporarily in volatile memory. Residual packet data can remain recoverable until the corresponding memory is reused, overwritten, deallocated, or otherwise altered.
The amount of network traffic recoverable from a memory image depends on several factors, including available memory, network utilization, operating-system networking behavior, buffer sizes, memory pressure, and the timing of the acquisition. A system experiencing heavy network activity may overwrite or recycle packet-buffer regions rapidly, reducing the amount of historical traffic that remains recoverable. Conversely, under some conditions, a lightly utilized system may retain older packet data in memory for a longer period. There is no fixed retention interval that can be assumed for a memory image.
The net scanner is available in Bulk Extractor and, depending on the version and configuration, can recover network packet structures and produce packets.pcap, a PCAP-formatted output file. This capability is particularly useful in memory forensics because residual network traffic may survive in volatile memory even when an original packet capture is unavailable. Several tools can be used to examine this output. Wireshark is well suited to this purpose. To open the carved output:
$ wireshark packets.pcap
The resulting packet listing should be treated as a carving result rather than a coherent network capture session. Examiners may encounter malformed or incomplete records, duplicate packets, missing packets, and traffic that cannot be placed reliably into its original sequence. Session-level reconstruction may therefore be impossible. For example, packets belonging to a TCP three-way handshake may be missing even though subsequent TCP traffic is recovered. This occurs because recovery depends on which portions of packet data and associated structures remained intact in memory when the image was acquired.
Examiners should also treat the timestamps displayed by Wireshark with caution. A carved packet does not necessarily retain the original capture-time metadata that a live packet-capture system associates with a packet. Consequently, the timestamps assigned to the generated PCAP may be zero, identical, synthetic, or otherwise unsuitable for establishing the original transmission time. Protocol-level timestamp fields, where present, are a separate matter and should not be confused with the PCAP capture timestamp. The PCAP timestamp should therefore not be used by itself to establish packet chronology or timing.
Note also that recovery of TCP-related structures in tcp.txt can be governed by version- and configuration-dependent scanner settings. If tcp.txt output is expected but absent, examiners should inspect the scanner configuration for the specific Bulk Extractor version being used rather than assuming that TCP output is always enabled by default.
Once the output is loaded, Wireshark display filters can be used to isolate traffic of evidentiary interest. The filter toolbar is located at the top of the window; filters such as http and dns provide useful starting points when those protocols are present in the recovered data. When examining DNS traffic, for example, the examiner can document recovered queries and corresponding responses. As with other memory-carved artifacts, however, these observations should be reported as recovered network evidence, not automatically as proof that a particular user initiated the communication.
The key principle is simple: a carved packet demonstrates that packet-related data was recoverable from the examined memory image; it does not, by itself, establish when the communication occurred, who initiated it, or that the resulting packet set represents a complete network session.
MAC addresses serve as globally unique identifiers for network interface cards. They are assigned by the manufacturer and, under normal circumstances, cannot be altered by the end user. Each address encodes both the Organizationally Unique Identifier (OUI) of the manufacturer and a device-specific serial number. The ability to uniquely attribute the source of network traffic is frequently of critical investigative value.
Examination of the file ether.txt reveals the full set of MAC addresses recovered from the memory image. Because of the volume of entries, analysis is more efficiently performed against the corresponding histogram file, ether_histogram.txt.
This is ether_histogram.txt, the histogram of Ethernet addresses recovered by the ether feature recorder (scan_net) from AsyncRAT_Win10.mem. It has seven entries, and several need interpretation before they support any conclusion.
Five entries fall in VMware-registered ranges: 00:0C:29:2D:0D:F3 (n=390), 00:0c:29:2d:0d:f3 (n=3), 00:0C:29:A6:0A:42 (n=275), 00:0C:29:BE:F2:C5 (n=115), and 00:50:56:FA:5C:A8 (n=1). These are four distinct addresses, since the first two are the same address in different letter cases. Together with the other evidence, this indicates the image came from a VMware guest, consistent with an analysis sandbox. 00:0C:29 is VMware's prefix for per-VM generated addresses. I believe 00:50:56:C0–FF is the range Workstation assigns to its own virtual network devices, which would make 00:50:56:FA:5C:A8 more likely the virtual gateway or NAT/DHCP service than a guest NIC.
Consider 3C:58:20:52:41:53 (n=3) and 68:99:20:52:41:53 (n=2). The last three bytes of each, 52 41 53, are ASCII "RAS". The first address decodes as <X RAS and the second as h, byte 0x99, space, RAS. A text string, plausibly a reference to the Remote Access Service, was read as six raw bytes that happened to pass the scanner's plausibility checks. Don't treat these as hardware addresses without corroboration, for example by checking whether the surrounding bytes look like a real Ethernet header followed by a valid IP header.
Does this file help with C2? More than a blanket "link-layer addresses don't cross routers" suggests. The TCP histogram places the C2 endpoint (192.168.135.20) and the subject (192.168.135.55) on what appears to be the same subnet, so the C2 host's MAC, and those of the DNS server and gateway, would appear in this file. The most frequent address is the natural candidate for the subject's own NIC, since it appears in every frame the guest sent or received in the recovered packets. To tie addresses to hosts, check the context around the ether features, ARP traffic in packets.pcap (Wireshark filter arp, then eth.addr == <mac>), and ip.txt entries at neighboring offsets. Without that mapping, this file only confirms the virtual environment. With it, you have the link-layer identity of the C2 host. It still says nothing about hosts off the local segment.
MAC addresses can support geolocation when they are Wi-Fi access point BSSIDs. A BSSID recovered from a device can be looked up in a wardriving database such as WiGLE to place the device near that access point. This applies only to genuine BSSIDs of physical wireless infrastructure that appear in such a database. None of the addresses here qualify, since they are virtual NIC addresses, so the technique does not apply to this image.
One of the options available in Bulk Extractor’s network scanner is the TCP histogram report. The resulting TCP histogram provides a frequency-ordered summary of all UDP and TCP packets extracted from the target data. In the present example, the highest-volume (“top talker”) endpoints are the subject system at 192.168.135.55 and a remote host at 192.168.135.20. The inclusion of port numbers permits the examiner to determine that the observed communication occurred over high ephemeral ports. It should be noted that false positives are possible in this data set, and the histogram does not constitute a complete enumeration of all network connections present in the image.
- The most frequent connections involve 192.168.135.20:8808 communicating with the subject host 192.168.135.55 on high ephemeral ports. Given that this is an AsyncRAT memory image, this traffic is highly suspicious and consistent with command-and-control (C2) or beaconing activity. Port 8808 is one of the three classic default ports used by AsyncRAT (the others being 6606 and 7707). Many real-world AsyncRAT samples and C2 servers are observed listening on or connecting to port 8808.
- Multiple connections to Microsoft Azure IP ranges on port 443 are present and likely represent legitimate or secondary cloud service traffic.
- Local DNS queries to 192.168.135.200 and a DHCP broadcast are consistent with normal internal network activity.
- The single garbled IPv6-style entry is likely an extraction artifact rather than a valid connection.
These sockets can be further correlated with Volatility’s netscan / netstat output and process memory for definitive attribution.
To correlate keywords identified by Bulk Extractor or a standalone strings utility with a process or kernel address space, the examiner may need to establish a physical-to-virtual mapping. This requires examining the relevant paging structures to determine whether the physical page containing the recovered bytes is mapped into a process or kernel virtual address space. The task can be complicated by shared pages, kernel memory, paged and non-paged pool allocations, memory that has been freed, and pages for which no current virtual mapping remains. A physical page may also be mapped into more than one address space. Consequently, the objective is not always to identify a single “owner,” but rather to determine which virtual address spaces, if any, currently map the physical location containing the artifact.
Both Rekall and Volatility provide mechanisms for performing this type of correlation. In Rekall, the vtop command can be used for virtual-to-physical translation. In Volatility 2.x, the legacy strings plugin can accept physical offsets and attempt to associate them with virtual address spaces; volshell also provides vtop functionality for address translation.
The Volatility 2 strings plugin accepts a formatted input file through the -s option. The input associates a physical-memory offset with the recovered string. The plugin then attempts to determine whether that physical location is mapped into a process or kernel address space and, where a mapping can be established, reports the corresponding virtual address. If no current mapping can be established — for example, because the page is no longer mapped — the result should be interpreted accordingly rather than as evidence of a particular process association.
A representative Volatility 2 invocation, given an input file named strings_pk.txt containing:
6382771058:pokestop
is:
$ vol.py -f dump.raw --profile=Win8SP1x64 strings -s strings_pk.txt 6382771058 [2340:0x0d69cb72] pokestop
The result indicates that the occurrence of pokestop was found at physical offset 6382771058 and that the physical location could be associated with virtual address 0x0d69cb72 in the address space associated with PID 2340. The examiner can then use pslist to resolve the PID:
$ vol.py -f dump.raw --profile=Win8SP1x64 pslist -p 2340 Offset(V) Name PID PPID Thds Hnds Sess Wow64 Start 0xffffe001e6920080 iexplore.exe 2340 4100 20 0 1 1 2016-08-03 03:27:59 UTC
This establishes that, in the captured memory state, the physical location containing the recovered occurrence was mapped into the virtual address space associated with iexplore.exe. That is considerably stronger evidence than merely observing the string at an unattributed physical-memory offset.
However, the correlation should not be overstated: the mapping establishes an association between the recovered bytes and the process's address space; it does not by itself prove that iexplore.exe created the string, that the user entered it, or that the process actively used the value. Additional corroborating artifacts are required to establish those conclusions. The examples above use Volatility 2, which is unmaintained and does not support recent Windows 10/11 builds. For current images, use Volatility 3. It needs no profile, relies on symbol tables, and provides windows.strings.Strings, which accepts GNU strings -t d output and maps each string to the processes whose address spaces contain it.
vol -f mem.raw windows.strings --strings-file ascii.txt
vol -f mem.raw windows.pslist --pid 2340In addition to analyzing memory dump files, examiners can analyze the Windows pagefile, typically pagefile.sys. The pagefile is part of Windows virtual-memory management and provides disk-backed storage for memory pages that the Memory Manager determines should no longer remain resident in physical RAM. It can therefore contain residual portions of process memory that are not otherwise available in the memory image.
It is important, however, not to treat the pagefile as a complete historical record of RAM. Pages may be discarded, reused, compressed, or remain backed by their original file rather than being written to the pagefile. Consequently, data recovered from pagefile.sys should be described as residual virtual-memory content recovered from the paging system, rather than simply as data that was once present in RAM.
A pagefile examined in isolation also provides limited process attribution. Windows does not ordinarily store a convenient process identifier alongside each pagefile page. An interesting string or other artifact recovered from a pagefile therefore cannot normally be attributed to a particular process merely from the pagefile offset itself.
Attribution can sometimes be improved when a memory image from the same acquisition is available. By examining the relevant page-table structures and page-table entries, including entries representing paged-out or transition states, an examiner may be able to correlate a pagefile-backed page with a virtual address space and consequently with a process. This is a more advanced PTE-level analysis and should not be confused with simple string searching of the pagefile.
Page-sized memory content can also span multiple pages. On the x86/x64 Windows systems commonly encountered in forensic examinations, the standard page size is 4 KB. Pages belonging to the same logical artifact are not necessarily adjacent in the pagefile, so recovery of larger artifacts can be incomplete or fragmented. The actual memory-management behavior should be considered when interpreting such results rather than assuming that every artifact occupies a contiguous region of the pagefile.
In some investigations, the existence of a recovered artifact may itself be significant even when process attribution cannot be established. For example, recovery of sensitive organizational data from a system that is not authorized to handle that data may provide an important investigative lead or indicate a potential policy or compliance concern. Whether the artifact constitutes an actual violation depends on the applicable policy, authorization, and regulatory context.
Modern versions of Windows also use memory compression as part of their memory-management strategy. Windows can compress infrequently accessed pages rather than immediately writing them to disk, with compressed data maintained in a system-managed compression store. Consequently, not all memory that has been removed from ordinary process working sets will necessarily appear in pagefile.sys. Compression behavior also means that the relationship between memory contents, the compression store, and pagefile contents is more complex than a simple “RAM → pagefile” model.
page_brute provides another approach for examining a pagefile or other large memory-related file. It divides the input into page-sized chunks (4 KB by default, where applicable) and applies YARA rules to those chunks. This can quickly identify chunks containing patterns associated with potentially relevant artifacts without requiring the examiner to inspect the entire pagefile manually.
The examiner can also supply custom YARA rules. The exact default ruleset is version-dependent, so the rule names should be verified against the version installed in the examination environment. For the version used in this post, examples include:
-
administrative_share_abuse -
remote_system_syntax -
httprequestheader -
httpresponseheader -
webartifact_html -
webartifact_javascript -
cmdshell -
social_security_syntax -
smtp_fragments -
irc -
webartifact_gmail -
ftp
A YARA match should be treated as a candidate artifact requiring contextual examination, not as proof that the associated application, protocol, or user activity occurred. The examiner should inspect the surrounding bytes and, where possible, corroborate the finding against the memory image, filesystem artifacts, event logs, process information, or other independent evidence.
The page_brute tool, developed by Michael Matonis (@matonis), is a Python utility designed to scan pagefiles and other large files for potentially relevant artifacts. It processes the input in 4 KB chunks by default and applies a set of YARA rules to each chunk.
Because the tool's default ruleset is resolved relative to its installation directory, the copy of page_brute maintained in the examiner's home directory should be executed from that directory. Change into the page_brute directory before running the tool. Alternatively, specify the YARA rules explicitly with the -r / --rules option. This behavior is version-dependent, so examiners should confirm the option names and ruleset location for the version installed in their examination environment.
You can run page_brute using its default signature set and provide an explicit name for the output directory with -o / --scanname. Providing a descriptive output name is preferable to relying on the tool's automatically generated name. When no output name is supplied, page_brute generates a directory name based on the execution timestamp, for example:
PAGE_BRUTE-2026-10-06-05:37:29-RESULTS
The automatically generated name does not identify the source pagefile or the YARA ruleset used for the scan. This can become confusing when multiple pagefiles are examined, or when the same pagefile is scanned repeatedly using different rule sets. For reproducible forensic work, use -o to assign a descriptive name that identifies the examination input and, where appropriate, the ruleset or scanning purpose.



.png)












Post a Comment