I've posted on the need for analysts to understand data structures before, and I recently conducted some analysis that I think illustrates this point extremely well.
The purpose of this blog post is to illustrate how certain artifacts can be used to detect the presence of malware on a system. While a tool for doing so is described, this is not a blog post about parsing histories for all of the browsers a user may or could have used, in part because the artifacts examined do not pertain to other browsers.
Not long ago, I tweeted that I'd written a plugin for the Forensic Scanner that gets statistics from IE (version 5 - 9) index.dat browser history files for all user profiles on the system. Almost immediately, someone tweeted asking, "what if the user isn't using IE?" That's a good question, but it misses the point of the analysis technique and of having the plugin in the first place.
I was analyzing a system recently that had been infected with ZeroAccess (see the Sophos report), and one of the things I was aware of the malware was capable of doing was click-fraud. In my analysis, I saw that the malware used an autostart persistence mechanism that was outside the scope of the user context...my timeline illustrated the artifacts being created. Knowing that much of the malware that communicates off-system will use the WinInet API functions to do so, I began looking at the index.dat files for the various user profiles available on the system. What I found was that the NetworkService account had much more significant "browser history" than the 'normal' user account on the system.
That's exactly right...there's no typo. The NetworkService account. How could that be? That's not something you see very often, is it? I mean, how does someone sit down at the keyboard and log into the account, and launch IE? The answer is...they don't. What happens is that when code using the WinInet API is run at privileges other than those of a user, the artifacts are created in another account profile. For example, back when Windows XP was more prevalent in my analysis lab, I would see systems on which the Default User profile had a populated index.dat file. I've seen the same thing with the LocalService account; this may depend upon which process the malware is injected into, and where that process falls in the svchost.exe hierarchy.
So my point is that for malware detection, checking all user accounts for statistics regarding their index.dat files might be a good idea. Once you understand the data structures in question - that is, the headers of the index.dat file, which, thanks to Joachim Metz, are well documented - this becomes a trivial task.
I started by writing a simple script that would parse the contents of the header of the index.dat file and tell me a little bit about what I could expect to see. Based on the format specification for the file, I was interested in things like the offset to the HASH table, as well as the directories beneath the "Temporary Internet Files\Content.IE5" folder and the number of cache files in each folder. This information is stored in the headers of the files, and is very easy to parse out and display. I got the script working and it proved to be very useful. However, I know that there's a process to using the script...I have to determine which user profiles are available, determine the version of Windows being examined, and based on those two pieces of information, type in the appropriate path to the index.dat file in question. By hand. Seriously?
So, I created a system class plugin for the Forensic Scanner to do all of this for me. Automatically. System class plugins are run against the entire system, whereas user class plugins are run against each user profile selected by the analyst. Based on the specific artifacts that I'm looking for, a system class plugin is exactly what I need.
What follows is an excerpt of the output from the ie_stats.pl plugin. First, the Administrator account profile:
g:\Documents and Settings\Administrator\Local Settings\Temporary Internet Files\Content.IE5\index.dat
File size : 163840
Hash Table Offset : 0x5000
Number of blocks : 1152
Number of alloc. blocks: 1089
Dir: LPVS8JVQ Files: 90
Dir: IR1PLUTE Files: 89
Dir: 5K3JMTA3 Files: 88
Dir: O3VB95DY Files: 89
As you can see, the Administrator account has some browser history associated with it. The hash table is located at offset 0x5000 within the index.dat file, and there are four subdirectories, each containing a number of cache files.
g:\Documents and Settings\Default User\Local Settings\Temporary Internet Files\Content.IE5\index.dat
File size : 32768
Hash Table Offset : 0x0
Number of blocks : 128
Number of alloc. blocks: 32
Dir: O8WMK2SC Files: 0
Dir: UAAUTN4C Files: 0
Dir: Q323MBYB Files: 0
Dir: 4VRPMD81 Files: 0
Okay, so this is what an empty index.dat looks like; the Default User profile has no IE history associated with it...there is no hash table, and the subdirectories don't contain any files.
g:\Documents and Settings\NetworkService\Local Settings\Temporary Internet Files\Content.IE5\index.dat
File size : 9437184
Hash Table Offset : 0x5000
Number of blocks : 73600
Number of alloc. blocks: 58784
Dir: EIQLTWH3 Files: 384
Dir: KFSPU8SK Files: 384
Dir: ABR75H1M Files: 384
Dir: 5QI7EWW7 Files: 384
Dir: TTYX5IX2 Files: 384
Dir: 5GDCH3XG Files: 383
Dir: KZZXKXWH Files: 383
Dir: LA2SZ5HL Files: 383
Dir: TLRE11UL Files: 383
Dir: IYT9OTGD Files: 383
Dir: 58I9EM25 Files: 383
....
The NetworkService account was the one that set off alarms! As you can see, this profile has a significant history! Parsing the actual index.dat file and displaying the entries in a micro-timeline (just the URL records) illustrated a significant amount of activity in a relatively short amount of time, with more URLs being requested per second than most users are capable of typing in or clicking via the browser.
In this case, I used the command line version of the Forensic Scanner to run a single plugin. I had mounted the image file via FTK Imager, and it appeared on my system as the G:\ volume. I then typed the following command:
C:\Perl\scanner>fscan.pl -s g:\windows\system32 -p ie_stats
That's all it took. Again, this is NOT a comprehensive analysis...this is a quick check to see if I could expect any potential issues. Another way to run this would be to select the "malware" artifact category...the ie_stats.pl plugin is included in that category. Also, the purpose of running this plugin is NOT to find indications of browser activity, for all browsers, and for all users. Instead, the purpose of this plugin is to check for something very specific...it does so automatically, accurately, and very, very quickly. This entire exercise took only a couple of minutes, most of which was spent mounting the image file...once that was done, the plugin ran very quickly, and provided me with the information I needed. I don't want to list all of the URL and REDR records from the NetworkService profile index.dat because just from what we see above (not all of the directories are listed) there are around 4000 or so files in the cache subdirectories.
Again, the purpose of this blog post is to illustrate an analysis technique. In the past, what I've done is to use ProDiscover to populate the "Internet History View" from within the image, and look for indications of service accounts with URL records in their index.dat. However, the methodology used by ProDiscover is more comprehensive...it searches the entire file system, and parses all of the records out of each index.dat file that it finds. In this case, that's much more than what I'm looking for.
This analysis technique can be combined with other tools into a more comprehensive process, as described in this blog post. Parsing application prefetch (*.pf) files and finding indications of wininet.dll as one of the loaded modules might be something to correlate with this analysis technique.
Resources
Rob Hensing's post on the Default User with an IE browser history
ForensicsWiki page: IE History File Format
Extremely relevant post from Hogfly (2007)
The Windows Incident Response Blog is dedicated to the myriad information surrounding and inherent to the topics of IR and digital analysis of Windows systems. This blog provides information in support of my books; "Windows Forensic Analysis" (1st thru 4th editions), "Windows Registry Forensics", as well as the book I co-authored with Cory Altheide, "Digital Forensics with Open Source Tools".
Saturday, March 16, 2013
Thursday, March 07, 2013
Wow6432Node: Registry Redirection
What is Registry redirection?
MS has a nice little explanation of Registry redirection and reflection here. Note that this MS page indicates that Registry reflection was removed as of Windows 7 and Windows 2008 R2.
In short, when a 32-bit application makes a call to write to the Registry on a 64-bit Windows system, by default, it doesn't go where we expect. A number of years ago, I was performing analysis of a 64-bit Windows 2003 server that we thought had been compromised via SQL injection...only I couldn't find the instance of the MS SQL Server listed in the Software hive. It turns out that I had to look beneath the Wow6432Node key
This page at MS provides information regarding Registry keys affected by Wow64.
What does it mean to us?
Well, for one...this is huge. No, I mean, it's HUGE. H. U. G. E. Why is that? Well, most of the malware I've seen over the years has been compiled for 32-bit platforms. While I haven't seen many 64-bit XP and 2003 systems (I have seen a few), I have seen a number of 64-bit Windows 7 systems, and all of the Windows 2008 R2 systems I've analyzed have been 64-bit.
So, if you're looking in the usual locations for malware...say, the Software\Microsoft\Windows\CurrentVersion\Run key, in either the HKLM or HKCU hives...then you're only going halfway and potentially missing a great deal of critical data in the Software\Wow6432Node\Microsoft\Windows\CurrentVersion\Run key. And unfortunately, most of us seem to be only going halfway.
This should be nothing new to the DFIR community. Redirection and virtualization of the Registry were discussed on pp 246 and 247, respectively, in Windows Forensic Analysis 2/e (published in 2009), as well as on pg 132 of WFAT 3/e. This topic is also discussed in this blog post.
Elizabeth S., from Google, discussed the Run keys extensively in her presentation at the 2012 SANS Forensic Summit. You'll notice that if you read through the presentation, she culled a lot of data about Registry keys from an AV vendor site. It appears that most of the testing platforms used by the vendor may be 32-bit, which may be why the Wow6432Node key isn't mentioned.
A good number of the RegRipper plugins that are affected by Registry redirection have been (or are being) updated to support Wow6432Node (where applicable, of course), and Corey Harrell has identified several others that need to be updated, as well. We're working on getting updated plugins into a new public distribution, so please bear with us.
Note that similar differences apply to the file system, as well, due to redirection. For example, on 64-bit systems, 32-bit applications can be found in the C:\Program Files (x86) and C:\Windows\SysWOW64 folders.
Resources
SamLogic.net article
Windows Confidential article
MS has a nice little explanation of Registry redirection and reflection here. Note that this MS page indicates that Registry reflection was removed as of Windows 7 and Windows 2008 R2.
In short, when a 32-bit application makes a call to write to the Registry on a 64-bit Windows system, by default, it doesn't go where we expect. A number of years ago, I was performing analysis of a 64-bit Windows 2003 server that we thought had been compromised via SQL injection...only I couldn't find the instance of the MS SQL Server listed in the Software hive. It turns out that I had to look beneath the Wow6432Node key
This page at MS provides information regarding Registry keys affected by Wow64.
What does it mean to us?
Well, for one...this is huge. No, I mean, it's HUGE. H. U. G. E. Why is that? Well, most of the malware I've seen over the years has been compiled for 32-bit platforms. While I haven't seen many 64-bit XP and 2003 systems (I have seen a few), I have seen a number of 64-bit Windows 7 systems, and all of the Windows 2008 R2 systems I've analyzed have been 64-bit.
So, if you're looking in the usual locations for malware...say, the Software\Microsoft\Windows\CurrentVersion\Run key, in either the HKLM or HKCU hives...then you're only going halfway and potentially missing a great deal of critical data in the Software\Wow6432Node\Microsoft\Windows\CurrentVersion\Run key. And unfortunately, most of us seem to be only going halfway.
This should be nothing new to the DFIR community. Redirection and virtualization of the Registry were discussed on pp 246 and 247, respectively, in Windows Forensic Analysis 2/e (published in 2009), as well as on pg 132 of WFAT 3/e. This topic is also discussed in this blog post.
Elizabeth S., from Google, discussed the Run keys extensively in her presentation at the 2012 SANS Forensic Summit. You'll notice that if you read through the presentation, she culled a lot of data about Registry keys from an AV vendor site. It appears that most of the testing platforms used by the vendor may be 32-bit, which may be why the Wow6432Node key isn't mentioned.
A good number of the RegRipper plugins that are affected by Registry redirection have been (or are being) updated to support Wow6432Node (where applicable, of course), and Corey Harrell has identified several others that need to be updated, as well. We're working on getting updated plugins into a new public distribution, so please bear with us.
Note that similar differences apply to the file system, as well, due to redirection. For example, on 64-bit systems, 32-bit applications can be found in the C:\Program Files (x86) and C:\Windows\SysWOW64 folders.
Resources
SamLogic.net article
Windows Confidential article
Tuesday, February 19, 2013
BinMode: Understanding Data Structures
As most analysts are aware, the tools we use provide a layer of abstraction over the data with which we're engaged. What we see can often depend upon the tool that we're using. For example, if a tool is written by a developer and the intended user is an administrator, then while the tool may be useful to a DFIR analyst, it may not provide all of the information that is truly useful to that analyst, based on the goals of their examination, the data that's actually presented by the tool, etc.
This is why understanding the data structures that we're working with can often be very beneficial.
By understanding what is actually available in the data structures, we can:
1. Make better use of the information that is available.
2. Locate deleted data in unallocated space (or other unstructured data)
A good recent example of this is the discussion of Java *.idx files, and the resulting parsers that have been created. Understanding the actual data structures that make up the headers and subsequent sections of these files lets us understand what we're looking at. For example, for a successful download, the header contains a field that tells us the size of the content. Most legitimate downloads also include this information in the server response, but malicious downloads of Java content don't always include this information. As such, we have not only have a good way for determining what may be a suspicious download, but we also have a pivot point we can use...we can use the content size to look for files of that size that were created on the system.
Another example of this is the IE history file format (thanks to Joachim for all the work he's done in documenting the format). A lot of analysts run various tools to parse out the user's IE web browser history, but how many understand what's actually in the structure? I'm not saying that you've memorized it and parse everything by hand, but rather that you know enough about it at least be curious when something is missing. For example, according to Joachim's documentation, each "URL" record can contain multiple time stamps, including when the page requested was last modified, last sync'd, and when it expires. I say "can" because in some cases, these may be set to 0. Further, according to the documentation, there's a flag setting that we can use to determine if the HTTP request was a GET or POST request.
How else can this be helpful? Mandiant's recently released APT intel report provides a great deal of useful information, including references to "wininet" on page 31. The WinInet API is what produces the artifacts most commonly associated with IE. As such, if your organization uses Firefox or Chrome, rather than IE, or you see that the "Default User", "NetworkService", or "LocalService" profiles begin developing quite large IE histories, this may be an indicator of activity.
I've used the documented format for Windows XP and 2003 Event Log records to not only parse the Event Log files on a system, but also locate and recover deleted event records from unallocated space. In fact, I had an instance where the intruder had cleared the Security Event Log after gaining access to the system, but I was able to recover 334 deleted event records from unallocated space, including the record that showed when they'd initially logged into the system.
Addendum, 20 Feb: I've mentioned DOSDate format time stamps in this blog before, as well as the data structures in which they're used (shell items in shellbag artifacts, ComDlg32 Registry subkey values, LNK files, Jump Lists, etc.). They're very pervasive across all Windows platforms, and more so on Windows 7 systems. This MS link provides some information regarding how these time stamps are constructed, as well as how they can play havok with timeline analysis if you're not familiar with them.
This doesn't apply just to Windows systems. A great example of this is Mari's recent blog post, Finding and Reverse Engineering Deleted SMS Messages. Mari provides a complete walk-thru of the data source being examined, going so far as to not only identify the data structure, but to also demonstrate how she did this, as well as to show how someone could go about identifying deleted SMS messages.
What's really interesting is that back in the day, there used to be more of a focus on understanding data structures. For example, some DF training programs would require candidates to parse partition tables and compute NTFS MFT data runs.
This is why understanding the data structures that we're working with can often be very beneficial.
By understanding what is actually available in the data structures, we can:
1. Make better use of the information that is available.
2. Locate deleted data in unallocated space (or other unstructured data)
A good recent example of this is the discussion of Java *.idx files, and the resulting parsers that have been created. Understanding the actual data structures that make up the headers and subsequent sections of these files lets us understand what we're looking at. For example, for a successful download, the header contains a field that tells us the size of the content. Most legitimate downloads also include this information in the server response, but malicious downloads of Java content don't always include this information. As such, we have not only have a good way for determining what may be a suspicious download, but we also have a pivot point we can use...we can use the content size to look for files of that size that were created on the system.
Another example of this is the IE history file format (thanks to Joachim for all the work he's done in documenting the format). A lot of analysts run various tools to parse out the user's IE web browser history, but how many understand what's actually in the structure? I'm not saying that you've memorized it and parse everything by hand, but rather that you know enough about it at least be curious when something is missing. For example, according to Joachim's documentation, each "URL" record can contain multiple time stamps, including when the page requested was last modified, last sync'd, and when it expires. I say "can" because in some cases, these may be set to 0. Further, according to the documentation, there's a flag setting that we can use to determine if the HTTP request was a GET or POST request.
How else can this be helpful? Mandiant's recently released APT intel report provides a great deal of useful information, including references to "wininet" on page 31. The WinInet API is what produces the artifacts most commonly associated with IE. As such, if your organization uses Firefox or Chrome, rather than IE, or you see that the "Default User", "NetworkService", or "LocalService" profiles begin developing quite large IE histories, this may be an indicator of activity.
I've used the documented format for Windows XP and 2003 Event Log records to not only parse the Event Log files on a system, but also locate and recover deleted event records from unallocated space. In fact, I had an instance where the intruder had cleared the Security Event Log after gaining access to the system, but I was able to recover 334 deleted event records from unallocated space, including the record that showed when they'd initially logged into the system.
Addendum, 20 Feb: I've mentioned DOSDate format time stamps in this blog before, as well as the data structures in which they're used (shell items in shellbag artifacts, ComDlg32 Registry subkey values, LNK files, Jump Lists, etc.). They're very pervasive across all Windows platforms, and more so on Windows 7 systems. This MS link provides some information regarding how these time stamps are constructed, as well as how they can play havok with timeline analysis if you're not familiar with them.
This doesn't apply just to Windows systems. A great example of this is Mari's recent blog post, Finding and Reverse Engineering Deleted SMS Messages. Mari provides a complete walk-thru of the data source being examined, going so far as to not only identify the data structure, but to also demonstrate how she did this, as well as to show how someone could go about identifying deleted SMS messages.
What's really interesting is that back in the day, there used to be more of a focus on understanding data structures. For example, some DF training programs would require candidates to parse partition tables and compute NTFS MFT data runs.
Training
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register. Each course includes access to tools and techniques that you won't find anywhere else, as well as a demonstration of the use of the Forensic Scanner.
On 10-12 June 2013, a Windows Forensic Analysis and Registry Analysis combo course will be hosted at the Santa Cruz PD training facility.
Course descriptions and other info on the courses is available here. Pricing for the combo course is $749 per seat, and will be listed on the ASI training page shortly.
Course descriptions and other info on the courses is available here. Pricing for the combo course is $749 per seat, and will be listed on the ASI training page shortly.
Saturday, February 16, 2013
Java, Timelines, and Training
A lot of interesting things have gone on lately...and I'm sure that there's more to come. I thought I'd take a moment to raise awareness of some of what happened recently,
Java In The News
Corey posted some links to tools the other day over on the jIIr blog. In the blog post, he talks about IDX parsers...not long ago, there was a flurry of activity around Java deployment cache index (*.idx) files, and what was contained in them.
You can find my IDX parsing code here. I updated it recently to support a variant to version 605 index files that I happened to see (that's one of the things I love about open source....).
Okay, so what's the big deal with this stuff? Is it just the "flavor-of-the-week" when it comes to DFIR analysis? I wouldn't think so...in fact, I'd suggest that if you do any sort of work with respect to malware or intrusion/compromise analysis, I would suggest to you that this is a pretty big deal. In fact, it's likely that anyone doing any sort of DFIR analysis is going to run up against this at some point...you may not recognize it at first, but it's likely to be there.
Remember this graphic from the SANS Forensic blog post from last year? Right in the middle of the super timeline, there are two *.idx files being created. This is from last year...during a recent examination, I added the metadata extracted from *.idx files to my own timeline, because the infection routine started with Java being run. I discuss that a little bit below.
Did anyone see this article recently? Java exploit? Yeah, apparently, that's what it was. I'd bet that the systems that were examined had some pretty interesting *.idx files on them.
I would suspect that as awareness of these artifacts are raised, analysts will begin to recognize that it's now easier to perform a root cause analysis, and to determine the initial infection vector. IMHO, the primary reason why this isn't done more often is because most assume that it takes too long to do an investigation; however, in not performing these investigations, we're missing out on a great deal of value information and intelligence that we can use to better protect our infrastructures.
Timeline Analysis
As I mentioned above, I used my own *.idx parser to add information from these files to a timeline...and this really helped my analysis.
First, as I do sometimes, I parsed the data separately and created a mini-timeline from just the data from the *.idx files. I do this, because in a full timeline, any times that data just gets lost. Now, I could have done something like created the events file, and then pulled out the information for the *.idx files using the type command, and piping that through find. Either way works.
Doing this showed me quite a bit about what the user likes to do during the day. I think what was most revealing about this data was that it made clear how much Java is used on the Internet. In a way, it was kind of fascinating, but at the same time, this didn't have much of anything to do with analysis goals. The fact was that while my timeline showed modifications to several cache directories when the system was being infected, what I did not see was the *.idx files that should have been part of the infection process.
One of the things I teach in my courses is that we need to understand artifacts so that you can not only recognize what's there, but so that you can also see when something that should be there isn't...and I had just such a case recently. There were a couple of possibilities as to what I was seeing, and one thing I found was that the installed AV had detected the malware; it quarantined the malware (i.e., created a file in the Quarantine directory) but the AV logs also stated quite clearly that (a) it was unable to remove the files from their original location, and (b) the AV software could not alert the user (notifications were disabled). Unfortunately, the logs did not go back all the way to the date of the infection, so I could not determine if the AV had actually detected the malware when it was installed, or if the detection I was seeing in the logs was a result of a product update that occurred after the infection. However, I did see that a file with the extension ".idx" had been created in the Quarantine folder, and had since been deleted. That wasn't definitive, but it was possible that the AV product had found something malicious in at least one *.idx file and had quarantined it. This would definitely account for the modification that occurred to the cache folder during the infection routine.
So, from the data I was looking at, there were two possibilities...one was that the AV had detected the infection, and while it was able to remove some of the files, it was unable to remove the actual malware itself...and since most AV products do nothing about the malware persistence mechanisms, the malware remained active (as confirmed by other artifacts).
The other possibility was that during the infection process, the malware "cleaned up" behind itself. After all, right there in the timeline, I could see that during the infection process, the Security Center was disabled, and the malware disabled the mechanism it used to gain access to the system. This was a multiple-stage infection...it started with Java via IE, downloaded something that ran and created files on the system. The second stage .exe file was deleted, but there were plenty of artifacts on the system that clearly showed not only that it had run, but also under what user context it had run. And it was clear that the infection process had also closed off the first stage of the infection.
So, there were some interesting findings in the analysis...all of the above was determined within the first 6 hours of analysis. As analysts, however, we're finders of fact, facts that we base and the data that I did have available had me leaning more toward the first possibility than the second. But the fact is that using these techniques, I had enough data to clearly identify not just the facts, but the gaps, as well.
Addendum: After my initial post, I saw this post at the KrebsOnSecurity blog, and this one on disabling Java for the browser on the Sophos site. Within minutes, I had written a RegRipper plugin to extract the UseJava2IExplorer value.
Resources
One of my own blog posts on the topic
ForensicsWiki Java Page
Mark Woan's format specification document
Training
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register. Each course includes access to tools and techniques that you won't find anywhere else, as well as a demonstration of the use of the Forensic Scanner.
On 10-12 June 2013, a Windows Forensic Analysis and Registry Analysis combo course will be hosted at the Santa Cruz PD training facility.
Course descriptions and other info on the courses is available here. Pricing for the combo course is $749 per seat, and will be listed on the ASI training page shortly.
Course descriptions and other info on the courses is available here. Pricing for the combo course is $749 per seat, and will be listed on the ASI training page shortly.
Wednesday, February 13, 2013
Hosted Training
On 10-12 June 2013, a Windows Forensic Analysis and Registry Analysis
combo course will be hosted at the Santa Cruz PD training facility.
Course descriptions and other info on the courses is available here. Pricing information for the combo course will be available on the ASI training page shortly.
If you are interested in registering for a seat in the training, please contact me here. As the date of the course approaches, information regarding parking, what to bring, etc., will be provided.
Course descriptions and other info on the courses is available here. Pricing information for the combo course will be available on the ASI training page shortly.
If you are interested in registering for a seat in the training, please contact me here. As the date of the course approaches, information regarding parking, what to bring, etc., will be provided.
Wednesday, February 06, 2013
There Are Four Lights: The Forensic Scanner
I made a push recently via social media to raise awareness about the Forensic Scanner, and based on some of what I saw come back, I'd like to take a moment to describe what the Forensic Scanner is, and perhaps clear us some misconceptions about the tool.
Just a quick reminder to everyone...in Nov, 2012, the Forensic Scanner moved from the Google Code site to this GitHub site. If you're going to try the Forensic Scanner, make sure that you run it as Administrator...if you have an admin account and you still have UAC enabled, you won't have what you think are full Admin rights on the box. Check out Corey's blog post on the topic.
First off, the Forensic Scanner is just a tool, nothing more. Like any other tool, if you don't understand how it was designed to be used, you very likely won't be using it to it's full capacity, or in it's most effective manner. Scanner applications have been used in various segments of infosec for quite some time. When I did vulnerability assessments back in the late '90s, we used scanner products to do some of the heavy lifting. Even today, there are scanners available for web app assessments, but the key point to remember is that these applications are not intended to replace analysts, or remove analysts from the picture. Instead, they are intended to perform a wide range of repeatable tests, so that an analyst can review the results and then focus their analysis efforts. This is also the intention of the Forensic Scanner.
The Forensic Scanner is:
A library of corporate knowledge/intel: An analyst may spend 8, 16, 24 hours or more in analysis and find something new. Take this finding, for example. One way to address a finding like this is for the analyst to keep it to himself...but that doesn't really help anyone, does it? An alternative might be to hold a brown-bag lunch with other analysts, put together a PPT, and tell them what you found. But how much more useful would it be to write a plugin, and share it with the other analysts? Within a few minutes, other analysts would have full access to the capability (i.e., finding the issue, or not...) without ever having to have the same experiences as the first analyst. On a team of eleven analysts, if it took 16 hours for the first analyst to find the issue, you've just saved the team, as a whole, 16 hrs x 10 analysts = 160 hours of time. This time can mean a great deal to your customer, as you will be providing with information they need to make critical business decisions in an extremely efficient manner.
By creating and maintaining plugins for these findings, the information is maintained in an accessible manner while the examiner who found the artifact is on vacation, or well after they left the organization. With the proper oversight, the plugin won't simply have lines of code...it will include references and explanations, so that the findings are not only repeatable, but they can be easily understood and explained.
How are you at memorizing the paths to various web browser history files on different versions of Windows (i.e., XP vs Windows 7)? How are you at mapping USB device usage? Do you want to have that available at the push of a button? That's what the Forensic Scanner can give you.
A force multiplier: By looking back at your last engagement and creating or updating plugins based on your findings, and then providing them to the team, you've bridged the gap between a checklist and actually implementing the checklist. This allows the experience of each analyst to be shared with others, which can lead to more work being done by the same number of analysts, in a much more efficient and timely manner.
A path to a competitive advantage: Analysts are going to find things that others either don't see, or haven't seen yet. As such, writing a plugin that you keep private within your team can lead to providing better, more comprehensive results to your customers, in a more timely manner. Based on the plugins you have in your library, you may be able to determine not only the malware that infected a system, but also determine the initial infection vector, in a much more timely manner. This means that you can provide not just findings, but intelligence to your customer, that they can then use to protect themselves.
The Forensic Scanner is NOT intended to replace any of the current tools that you own and use. Rather, the purpose of the Forensic Scanner is to augment and optimize your use of those tools that you already own, and get you to the point of deep engagement with those tools much sooner.
Deployment Scenarios
When I first had the idea for the Forensic Scanner,
Lab Tech: A lab tech receives an image, and as part of the verification and in-processing procedures, runs a scan of the mounted image. The lab tech then contacts the designated analyst to let her know that the image and report are in a specific location, either in the "cloud", or to be retrieved in some other manner. Rather than having to run through all of the checks herself, the examiner can review the report and focus her analysis faster, providing much more comprehensive and timely findings.
LE examiner: An LE examiner might be interested in P2P file sharing, and one of the biggest issues for LE (at all levels) is that the examiners are cops first. This makes it very difficult to keep up on various analysis techniques and artifacts, but the Forensic Scanner puts that right at your fingertips. Perhaps your cases involving illicit images don't require you to do much more than find and catalog the images, and you're done. Or perhaps you need more...did the user actually access the images at any point? Did the transfer of the files involve a USB storage device of any kind? Was a digital camera or a smartphone connected to the system?
Consultant: A consultant or even an IT security staff member can be on-site, performing triage and acquisitions. They can run a report, and because the report contains no sensitive (PII, PHI, PCI, etc.) information, they can archive/protect the report, and ship it off to another analyst who is off-site, who can then perform analysis of the report. The analyst can then respond to the on-site consultant, providing information that can then help them focus their efforts ("..acquire these 5 systems instead of all 300...").
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register. Each course includes access to tools and techniques that you won't find anywhere else, as well as a demonstration of the use of the Forensic Scanner.
Just a quick reminder to everyone...in Nov, 2012, the Forensic Scanner moved from the Google Code site to this GitHub site. If you're going to try the Forensic Scanner, make sure that you run it as Administrator...if you have an admin account and you still have UAC enabled, you won't have what you think are full Admin rights on the box. Check out Corey's blog post on the topic.
First off, the Forensic Scanner is just a tool, nothing more. Like any other tool, if you don't understand how it was designed to be used, you very likely won't be using it to it's full capacity, or in it's most effective manner. Scanner applications have been used in various segments of infosec for quite some time. When I did vulnerability assessments back in the late '90s, we used scanner products to do some of the heavy lifting. Even today, there are scanners available for web app assessments, but the key point to remember is that these applications are not intended to replace analysts, or remove analysts from the picture. Instead, they are intended to perform a wide range of repeatable tests, so that an analyst can review the results and then focus their analysis efforts. This is also the intention of the Forensic Scanner.
The Forensic Scanner is:
A library of corporate knowledge/intel: An analyst may spend 8, 16, 24 hours or more in analysis and find something new. Take this finding, for example. One way to address a finding like this is for the analyst to keep it to himself...but that doesn't really help anyone, does it? An alternative might be to hold a brown-bag lunch with other analysts, put together a PPT, and tell them what you found. But how much more useful would it be to write a plugin, and share it with the other analysts? Within a few minutes, other analysts would have full access to the capability (i.e., finding the issue, or not...) without ever having to have the same experiences as the first analyst. On a team of eleven analysts, if it took 16 hours for the first analyst to find the issue, you've just saved the team, as a whole, 16 hrs x 10 analysts = 160 hours of time. This time can mean a great deal to your customer, as you will be providing with information they need to make critical business decisions in an extremely efficient manner.
By creating and maintaining plugins for these findings, the information is maintained in an accessible manner while the examiner who found the artifact is on vacation, or well after they left the organization. With the proper oversight, the plugin won't simply have lines of code...it will include references and explanations, so that the findings are not only repeatable, but they can be easily understood and explained.
How are you at memorizing the paths to various web browser history files on different versions of Windows (i.e., XP vs Windows 7)? How are you at mapping USB device usage? Do you want to have that available at the push of a button? That's what the Forensic Scanner can give you.
A force multiplier: By looking back at your last engagement and creating or updating plugins based on your findings, and then providing them to the team, you've bridged the gap between a checklist and actually implementing the checklist. This allows the experience of each analyst to be shared with others, which can lead to more work being done by the same number of analysts, in a much more efficient and timely manner.
A path to a competitive advantage: Analysts are going to find things that others either don't see, or haven't seen yet. As such, writing a plugin that you keep private within your team can lead to providing better, more comprehensive results to your customers, in a more timely manner. Based on the plugins you have in your library, you may be able to determine not only the malware that infected a system, but also determine the initial infection vector, in a much more timely manner. This means that you can provide not just findings, but intelligence to your customer, that they can then use to protect themselves.
The Forensic Scanner is NOT intended to replace any of the current tools that you own and use. Rather, the purpose of the Forensic Scanner is to augment and optimize your use of those tools that you already own, and get you to the point of deep engagement with those tools much sooner.
Deployment Scenarios
When I first had the idea for the Forensic Scanner,
Lab Tech: A lab tech receives an image, and as part of the verification and in-processing procedures, runs a scan of the mounted image. The lab tech then contacts the designated analyst to let her know that the image and report are in a specific location, either in the "cloud", or to be retrieved in some other manner. Rather than having to run through all of the checks herself, the examiner can review the report and focus her analysis faster, providing much more comprehensive and timely findings.
LE examiner: An LE examiner might be interested in P2P file sharing, and one of the biggest issues for LE (at all levels) is that the examiners are cops first. This makes it very difficult to keep up on various analysis techniques and artifacts, but the Forensic Scanner puts that right at your fingertips. Perhaps your cases involving illicit images don't require you to do much more than find and catalog the images, and you're done. Or perhaps you need more...did the user actually access the images at any point? Did the transfer of the files involve a USB storage device of any kind? Was a digital camera or a smartphone connected to the system?
Consultant: A consultant or even an IT security staff member can be on-site, performing triage and acquisitions. They can run a report, and because the report contains no sensitive (PII, PHI, PCI, etc.) information, they can archive/protect the report, and ship it off to another analyst who is off-site, who can then perform analysis of the report. The analyst can then respond to the on-site consultant, providing information that can then help them focus their efforts ("..acquire these 5 systems instead of all 300...").
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register. Each course includes access to tools and techniques that you won't find anywhere else, as well as a demonstration of the use of the Forensic Scanner.
Saturday, February 02, 2013
BinMode: Parsing Java *.idx files, pt trios
Things have progressed a great deal since I last blogged on this subject. Specifically, additional information and resources have been added to the ForensicsWiki page on this topic, and Brian has updated his Python code. Mark Woan has created a .Net console application for parsing these files, as well, and his repo contains a PDF document that delineates the structure of the various versions of these files.
Running my own tool against the Java deployment cache on my system, I don't see much in the way of interesting data; most of what I have on this system is the result of accessing SANS webcasts. However, parsing the data from the *.idx file that Corey provided, we see the following:
File: d:\cases\781da39f-6b6c0267.idx
URL: http://xhaito.com/work/builds/exp_files/rox.jar
IP: 91.213.217.31
content-length: 14226
last-modified: Sun Sep 12 15:15:32 2010 UTC
Server Response:
------------------------------
HTTP/1.1 200 OK
content-length: 14226
last-modified: Sun, 12 Sep 2010 15:15:32 GMT
content-type: text/plain
date: Sun, 12 Sep 2010 22:38:35 GMT
server: Apache/2
deploy-request-content-type: application/x-java-archive
The information displayed at the top of the output, above "Server Response", is from the header of the *.idx file, while the rest of the information is from Section 2 of the file. For specifics of this data, take a look at the PDF document that Mark provided. Suffice to say, this is a great resource, because what you're seeing is extracted from the binary contents of the file. Yes, the strings for the URL and IP address can be found via a text or keyword search, but an understanding of the data source and the data structure provides valuable context to the search hits. Even better, a targeted, Sniper Forensics approach to going after the data is something that we can do now because of what we know about the data itself.
Okay...so what? Now that we have this information available, how do we use it in exams? Perhaps the most obvious would be to parse the contents of the *.idx files and check the output against the Malware Domain List, or "MDL".
Another method of analysis for this information would be to parse the data and correlate statistics from all of the available *.idx files (URL, IP address, content type, etc.), showing the stats as an overview before digging into the data itself. Combining that two...MDL check and stats...would be a great way to perform data reduction. One might incorporate checking against the MDL directly into a tool that parses the data within *.idx files for inclusion directly into a timeline, adding the pivot points directly to the timeline itself. Incorporating this with other data...specifically, the user's web browser history...would allow an analyst to easily 'see' an Initial Infection Vector.
For me, the first step is to incorporate this information into a timeline...
Addendum:I updated my code recently to provide more than the output that you see above. The new version includes options for CSV or TLN output. It also includes a heuristic to help detect potentially malicious Java archives, as opposed to those that may be legit.
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register.
Running my own tool against the Java deployment cache on my system, I don't see much in the way of interesting data; most of what I have on this system is the result of accessing SANS webcasts. However, parsing the data from the *.idx file that Corey provided, we see the following:
File: d:\cases\781da39f-6b6c0267.idx
URL: http://xhaito.com/work/builds/exp_files/rox.jar
IP: 91.213.217.31
content-length: 14226
last-modified: Sun Sep 12 15:15:32 2010 UTC
Server Response:
------------------------------
HTTP/1.1 200 OK
content-length: 14226
last-modified: Sun, 12 Sep 2010 15:15:32 GMT
content-type: text/plain
date: Sun, 12 Sep 2010 22:38:35 GMT
server: Apache/2
deploy-request-content-type: application/x-java-archive
The information displayed at the top of the output, above "Server Response", is from the header of the *.idx file, while the rest of the information is from Section 2 of the file. For specifics of this data, take a look at the PDF document that Mark provided. Suffice to say, this is a great resource, because what you're seeing is extracted from the binary contents of the file. Yes, the strings for the URL and IP address can be found via a text or keyword search, but an understanding of the data source and the data structure provides valuable context to the search hits. Even better, a targeted, Sniper Forensics approach to going after the data is something that we can do now because of what we know about the data itself.
Okay...so what? Now that we have this information available, how do we use it in exams? Perhaps the most obvious would be to parse the contents of the *.idx files and check the output against the Malware Domain List, or "MDL".
Another method of analysis for this information would be to parse the data and correlate statistics from all of the available *.idx files (URL, IP address, content type, etc.), showing the stats as an overview before digging into the data itself. Combining that two...MDL check and stats...would be a great way to perform data reduction. One might incorporate checking against the MDL directly into a tool that parses the data within *.idx files for inclusion directly into a timeline, adding the pivot points directly to the timeline itself. Incorporating this with other data...specifically, the user's web browser history...would allow an analyst to easily 'see' an Initial Infection Vector.
For me, the first step is to incorporate this information into a timeline...
Addendum:I updated my code recently to provide more than the output that you see above. The new version includes options for CSV or TLN output. It also includes a heuristic to help detect potentially malicious Java archives, as opposed to those that may be legit.
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register.
Monday, January 28, 2013
Are You Being Served, pt II
This article isn't going to be directed toward digital analysts; rather, it will be directed more to folks who hire or contract with analysts or firms, and are the recipients (or customers) of the technical work performed by those digital forensics analysts. My goal here is to simply express some thoughts on how customers might go about determining if the results of the work that they contracted for are meeting their needs.
Previously in this blog, I asked the question, Are you being served? If you've asked yourself that question, you may be wondering...how would I know? Selecting a DFIR analyst (either an individual or a firm) is really no different for evaluating and hiring any other provider of services, such as a plumber or auto mechanic. The difference is that plumbers and mechanics fix something for you, and you can evaluate their services based on if the problem is fixed, and for how long. For customers of digital analysis services, determining if you're getting what you paid for is a bit more difficult.
In exploring the subject of finding a digital forensics expert, I ran across this article at the Law.com web site. The article contains a number of aspects of the overall digital analysis services that lawyers should consider when looking for a digital forensics expert. For example, the article suggests that when asked to identify methods of data exfiltration, analysts should include USB devices. This is good to know, but more importantly, does the analyst identify all such devices, or only the thumb drives? How do you know? Does the analyst make an attempt to determine the use of counter-forensics techniques, where a user might delete certain artifacts in an attempt to hide the fact that they connected a specific device to the system? What details can the analyst provide with respect to the device being connected to the system, and how a user may have interacted with that device? Regardless of the data exfiltration method used (USB device, web mail, BlueTooth, etc.), how does the analyst address data movement, in particular?
Beyond those items addressed in the article, some other things to consider include (but are not limited to):
Does the analyst explore historical data, such as Volume Shadow Copies (VSCs), when and where it is appropriate to do so? If not, why? If the methodology used by the analyst fails to find any VSCs, what does the analyst state as the reason for this finding?
What about other artifacts? When the analyst provides a finding, do they have additional artifacts to support their findings, or are their findings based on that one artifact? If artifacts (such as Prefetch files) are not examined or missing, what reason does the analyst provide?
If you're interested in the existence of malware on a system, what does the analyst do to address this issue? Do they run AV against the mounted image? What else do they do? If malware is found, do they determine the initial infection vector? Do they determine if the malware ever actually executed?
When you look at the report provided, does the information in it answer your questions and address your concerns, or are there gaps? Does the analyst connect the dots in the report, or do they skip over many of the dots, and fill in the gaps using speculation?
One question that you might consider asking is, what tools does the analyst use, but I would suggest that it's more important to know how the tools are used. For example, having access to one of the commercial analysis suites can be a good thing, particularly if the analyst states that they will use it on your case to perform a keyword search. But does it make sense to do so? Did they work with you to develop a list of keywords to use in the search? I've heard of examinations that were delayed for some time while the data was being preprocessed and indexed in preparation for a keyword search, yet none of the analysts could state why the keyword search was necessary or of value to the case itself.
There is often much more to digital analysis than simply finding one or two artifacts in order to "solve the case". Systems today are sufficiently complex that multiple artifacts are needed to identify the context of a single artifact, such as a tool not finding VSCs within an image of a Windows 7 system. Digital analysis is very often used as the basis for making critical business decisions or addressing legal questions, so the question remains...are you being served? Are you getting the data that you need, in a timely manner, and in a manner that you can understand and use?
Resources
Law.com - How to Find a Digital Forensics Expert
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register.
Previously in this blog, I asked the question, Are you being served? If you've asked yourself that question, you may be wondering...how would I know? Selecting a DFIR analyst (either an individual or a firm) is really no different for evaluating and hiring any other provider of services, such as a plumber or auto mechanic. The difference is that plumbers and mechanics fix something for you, and you can evaluate their services based on if the problem is fixed, and for how long. For customers of digital analysis services, determining if you're getting what you paid for is a bit more difficult.
In exploring the subject of finding a digital forensics expert, I ran across this article at the Law.com web site. The article contains a number of aspects of the overall digital analysis services that lawyers should consider when looking for a digital forensics expert. For example, the article suggests that when asked to identify methods of data exfiltration, analysts should include USB devices. This is good to know, but more importantly, does the analyst identify all such devices, or only the thumb drives? How do you know? Does the analyst make an attempt to determine the use of counter-forensics techniques, where a user might delete certain artifacts in an attempt to hide the fact that they connected a specific device to the system? What details can the analyst provide with respect to the device being connected to the system, and how a user may have interacted with that device? Regardless of the data exfiltration method used (USB device, web mail, BlueTooth, etc.), how does the analyst address data movement, in particular?
Beyond those items addressed in the article, some other things to consider include (but are not limited to):
Does the analyst explore historical data, such as Volume Shadow Copies (VSCs), when and where it is appropriate to do so? If not, why? If the methodology used by the analyst fails to find any VSCs, what does the analyst state as the reason for this finding?
What about other artifacts? When the analyst provides a finding, do they have additional artifacts to support their findings, or are their findings based on that one artifact? If artifacts (such as Prefetch files) are not examined or missing, what reason does the analyst provide?
If you're interested in the existence of malware on a system, what does the analyst do to address this issue? Do they run AV against the mounted image? What else do they do? If malware is found, do they determine the initial infection vector? Do they determine if the malware ever actually executed?
When you look at the report provided, does the information in it answer your questions and address your concerns, or are there gaps? Does the analyst connect the dots in the report, or do they skip over many of the dots, and fill in the gaps using speculation?
One question that you might consider asking is, what tools does the analyst use, but I would suggest that it's more important to know how the tools are used. For example, having access to one of the commercial analysis suites can be a good thing, particularly if the analyst states that they will use it on your case to perform a keyword search. But does it make sense to do so? Did they work with you to develop a list of keywords to use in the search? I've heard of examinations that were delayed for some time while the data was being preprocessed and indexed in preparation for a keyword search, yet none of the analysts could state why the keyword search was necessary or of value to the case itself.
There is often much more to digital analysis than simply finding one or two artifacts in order to "solve the case". Systems today are sufficiently complex that multiple artifacts are needed to identify the context of a single artifact, such as a tool not finding VSCs within an image of a Windows 7 system. Digital analysis is very often used as the basis for making critical business decisions or addressing legal questions, so the question remains...are you being served? Are you getting the data that you need, in a timely manner, and in a manner that you can understand and use?
Resources
Law.com - How to Find a Digital Forensics Expert
Interested in Windows DFIR training? Windows Forensic Analysis, 11-12 Mar; Timeline Analysis, 9-10 Apr. Pricing and Calendar. Send email here to register.
Why "BinMode"?
You may be wondering why I've started posting articles to my blog with titles that start with "BinMode" and "There Are Four Lights".
The "BinMode" posts are dedicated to deeply technical posts; the name comes from the fact that sometimes I'll write a Perl script that requires me to open a file using binmode(), so that I can parse the file on a binary level. These are generally posts that go beyond the tools, which tend to provide a layer of abstraction between the data and analyst. I feel that it's important for analysts to understand what data is available to them, so that they can make better decisions as to which tool to use to extract and process that data.
An example of this is the recent work I've done parsing the Java deployment cache index (*.idx) files. Beyond opening these files in a hex editor, one resource that I had access to in order to assist me in parsing the files is this source code page: CacheEntry.java. Another resource that became available later in the process is the format specification that Mark Woan documented. What these resources show is that within the binary data, there is potentially some extremely valuable information. This information might be most useful during a root cause analysis investigation, perhaps to determine the initial infection vector of malware, or how a compromise occurred.
The "Four Lights" articles are partly a nod to the inner geek (and Star Trek fan) in all of us, but they're also to address something that may be lesser known, or perhaps seen as a misconception within the digital forensic analysis community. The title alludes to an episode of ST:TNG, during which his captors attempted to get the greatest starship captain...EVER...to say that there were only three lights, when, in fact, there were four.
If there is a particular topic that you'd like me to expand upon, or if there's something that you'd like to see addressed, feel free to leave a comment here, or to send me an email.
Interested in Windows DF training? Check it out: Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar. Be sure to check the WindowsIR Training Page for updates.
The "BinMode" posts are dedicated to deeply technical posts; the name comes from the fact that sometimes I'll write a Perl script that requires me to open a file using binmode(), so that I can parse the file on a binary level. These are generally posts that go beyond the tools, which tend to provide a layer of abstraction between the data and analyst. I feel that it's important for analysts to understand what data is available to them, so that they can make better decisions as to which tool to use to extract and process that data.
An example of this is the recent work I've done parsing the Java deployment cache index (*.idx) files. Beyond opening these files in a hex editor, one resource that I had access to in order to assist me in parsing the files is this source code page: CacheEntry.java. Another resource that became available later in the process is the format specification that Mark Woan documented. What these resources show is that within the binary data, there is potentially some extremely valuable information. This information might be most useful during a root cause analysis investigation, perhaps to determine the initial infection vector of malware, or how a compromise occurred.
The "Four Lights" articles are partly a nod to the inner geek (and Star Trek fan) in all of us, but they're also to address something that may be lesser known, or perhaps seen as a misconception within the digital forensic analysis community. The title alludes to an episode of ST:TNG, during which his captors attempted to get the greatest starship captain...EVER...to say that there were only three lights, when, in fact, there were four.
If there is a particular topic that you'd like me to expand upon, or if there's something that you'd like to see addressed, feel free to leave a comment here, or to send me an email.
Interested in Windows DF training? Check it out: Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar. Be sure to check the WindowsIR Training Page for updates.
Monday, January 21, 2013
BinMode: Parsing Java *.idx files, pt. deux
My last post addressed parsing Java *.idx files, and since I released that post, a couple of resources related to the post have been updated. In particular, Joachim Metz has updated the ForensicsWiki page he started to include more information about the format of the *.idx files, with some information specific to what is thought to be the header of the files.
Also, Corey Harrell was kind enough to share the *.idx file from this blog post with me (click here to see the graphic of what the file "looks like" in Corey's post), and I ran it through the parser to see what I could find:
File: d:\test\781da39f-6b6c0267.idx
Times from header:
------------------------------
time_0: Sun Sep 12 15:15:32 2010 UTC
time_2: Sun Sep 12 22:38:40 2010 UTC
URL: http://xhaito.com/work/builds/exp_files/rox.jar
IP: 91.213.217.31
Server Response:
------------------------------
HTTP/1.1 200 OK
content-length: 14226
last-modified: Sun, 12 Sep 2010 15:15:32 GMT
content-type: text/plain
date: Sun, 12 Sep 2010 22:38:35 GMT
server: Apache/2
deploy-request-content-type: application/x-java-archive
Ah, pretty interesting stuff. Again, the "Times from header" section is comprised of, at this moment, data from those offsets within the header that Joachim has identified as possibly being time stamps. In the code, I have it display only those times that are not zero. What we don't have at the moment is information about the structure of the header so that we can identify to what the time stamps refer.
However, this code can be used to parse *.idx files and help determine to what the times refer. For example, in the output above we see that "time_0" is equivalent to the "last modified" field in the server response, and that the "time_2" field is a few seconds after the "date" field in the server response. Perhaps incorporating this information into a timeline might be useful, while research continues in order to identify what the time stamps represent. What is very useful is that the *.idx files are associated with a specific user profile, so for testing purposes, an analyst should be able to incorporate browser history and *.idx info into a timeline, and perhaps be able to "see" what the time stamps may refer to...if the analyst were to control the entire test environment, to include the web server, even more information may be developed.
Speaking of timelines, Sploited commented to my previous post regarding developing timelines analysis pivot points from other resources; in the comment, a script for parsing IE history files (urlcache.pl) was mentioned; I would suggest that incorporating a user's web history, as well as incorporating searches against the Malware Domain List might be extremely helpful in identifying initial infect vectors and entry points.
Interested in Windows DF training? Check it out: Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar. Be sure to check the WindowsIR Training Page for updates.
Also, Corey Harrell was kind enough to share the *.idx file from this blog post with me (click here to see the graphic of what the file "looks like" in Corey's post), and I ran it through the parser to see what I could find:
File: d:\test\781da39f-6b6c0267.idx
Times from header:
------------------------------
time_0: Sun Sep 12 15:15:32 2010 UTC
time_2: Sun Sep 12 22:38:40 2010 UTC
URL: http://xhaito.com/work/builds/exp_files/rox.jar
IP: 91.213.217.31
Server Response:
------------------------------
HTTP/1.1 200 OK
content-length: 14226
last-modified: Sun, 12 Sep 2010 15:15:32 GMT
content-type: text/plain
date: Sun, 12 Sep 2010 22:38:35 GMT
server: Apache/2
deploy-request-content-type: application/x-java-archive
Ah, pretty interesting stuff. Again, the "Times from header" section is comprised of, at this moment, data from those offsets within the header that Joachim has identified as possibly being time stamps. In the code, I have it display only those times that are not zero. What we don't have at the moment is information about the structure of the header so that we can identify to what the time stamps refer.
However, this code can be used to parse *.idx files and help determine to what the times refer. For example, in the output above we see that "time_0" is equivalent to the "last modified" field in the server response, and that the "time_2" field is a few seconds after the "date" field in the server response. Perhaps incorporating this information into a timeline might be useful, while research continues in order to identify what the time stamps represent. What is very useful is that the *.idx files are associated with a specific user profile, so for testing purposes, an analyst should be able to incorporate browser history and *.idx info into a timeline, and perhaps be able to "see" what the time stamps may refer to...if the analyst were to control the entire test environment, to include the web server, even more information may be developed.
Speaking of timelines, Sploited commented to my previous post regarding developing timelines analysis pivot points from other resources; in the comment, a script for parsing IE history files (urlcache.pl) was mentioned; I would suggest that incorporating a user's web history, as well as incorporating searches against the Malware Domain List might be extremely helpful in identifying initial infect vectors and entry points.
Interested in Windows DF training? Check it out: Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar. Be sure to check the WindowsIR Training Page for updates.
Saturday, January 19, 2013
BinMode: Parsing Java *.idx files
One of the Windows artifacts that I talk about in my training courses is application log files, and I tend to sort of gloss over this topic, simply because there are so many different kinds of log files produced by applications. Some applications, in particular AV, will write their logs to the Application Event Log, as well as a text file. I find this to be very useful because the Application Event Log will "roll over" as it gathers more events; most often, the text logs will continue to be written to by the application. I talk about these logs in general because it's important for analysts to be aware of them, but I don't spend a great deal of time discussing them because we could be there all week talking about them.
With the recent (Jan, 2013) issues regarding a Java 0-day vulnerability, my interest in artifacts of compromise were piqued yet again when I found that someone had released some Python code for parsing Java deployment cache *.idx files. I located the *.idx files on my own system, opened a couple of them up in a hex editor and began conducting pattern analysis to see if I could identify a repeatable structure. I found enough information to create a pretty decent parser for the *.idx files to which I have access.
Okay, so the big question is...so what? Who cares? Well, Corey Harrell had an excellent post to his blog regarding Finding (the) Initial Infection Vector, which I think is something that folks don't do often enough. Using timeline analysis, Corey identified artifacts that required closer examination; using the right tools and techniques, this information can also be included directly into the timeline (see the Sploited blog post listed in the Resources section below) to provide more context to the timeline activity.
The testing I've been able to do with the code I wrote has been somewhat limited, as I haven't had a system that might be infected come across my desk in a bit, and I don't have access to an *.idx file like what Corey illustrated in his blog post (notice that it includes "pragma" and "cache control" statements). However, what I really like about the code is that I have access to the data itself, and I can modify the code to meet my analysis needs, much the way I did with the Prefetch file analysis code that I wrote. For example, I can perform frequency analysis of IP addresses or URLs, server types, etc. I can perform searches for various specific data elements, or simply run the output of the tool through the find command, just to see if something specific exists. Or, I can have the code output information in TLN format for inclusion in a timeline.
Regardless of what I do with the code itself, I know have automatic access to the data, and I have references included in the script itself; as such, the headers of the script serve as documentation, as well as a reminder of what's being examined, and why. This bridges the gap between having something I need to check listed in a spreadsheet, and actually checking or analyzing those artifacts.
Resources
ForensicsWiki Page: Java
Sploited blog post: Java Forensics Using TLN Timelines
jIIr: Almost Cooked Up Some Java, Finding Initial Infection Vector
Interested in Windows DF training? Check it out: Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar.
With the recent (Jan, 2013) issues regarding a Java 0-day vulnerability, my interest in artifacts of compromise were piqued yet again when I found that someone had released some Python code for parsing Java deployment cache *.idx files. I located the *.idx files on my own system, opened a couple of them up in a hex editor and began conducting pattern analysis to see if I could identify a repeatable structure. I found enough information to create a pretty decent parser for the *.idx files to which I have access.
Okay, so the big question is...so what? Who cares? Well, Corey Harrell had an excellent post to his blog regarding Finding (the) Initial Infection Vector, which I think is something that folks don't do often enough. Using timeline analysis, Corey identified artifacts that required closer examination; using the right tools and techniques, this information can also be included directly into the timeline (see the Sploited blog post listed in the Resources section below) to provide more context to the timeline activity.
The testing I've been able to do with the code I wrote has been somewhat limited, as I haven't had a system that might be infected come across my desk in a bit, and I don't have access to an *.idx file like what Corey illustrated in his blog post (notice that it includes "pragma" and "cache control" statements). However, what I really like about the code is that I have access to the data itself, and I can modify the code to meet my analysis needs, much the way I did with the Prefetch file analysis code that I wrote. For example, I can perform frequency analysis of IP addresses or URLs, server types, etc. I can perform searches for various specific data elements, or simply run the output of the tool through the find command, just to see if something specific exists. Or, I can have the code output information in TLN format for inclusion in a timeline.
Regardless of what I do with the code itself, I know have automatic access to the data, and I have references included in the script itself; as such, the headers of the script serve as documentation, as well as a reminder of what's being examined, and why. This bridges the gap between having something I need to check listed in a spreadsheet, and actually checking or analyzing those artifacts.
Resources
ForensicsWiki Page: Java
Sploited blog post: Java Forensics Using TLN Timelines
jIIr: Almost Cooked Up Some Java, Finding Initial Infection Vector
Interested in Windows DF training? Check it out: Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar.
Saturday, January 12, 2013
There Are Four Lights: The Analysis Matrix
I've talked a lot in this blog about employing event categories when developing, and in particular, when analyzing timelines, and the fact is that we can use these categories for much more that just adding analysis functionality to our timelines. In fact, using artifact and event categories can greatly enhance our overall analysis capabilities. This is something that Corey Harrell and I have spent a great deal of time discussing.
For one, if we categorize events, we can raise our level of awareness of the context of the data that we're analyzing. Having categories for various artifacts can help us increase our relative level of confidence in the data that we're analyzing, because instead of looking at just one artifact, we're going to be looking at various similar, related artifacts together.
Another benefit of artifact categories is that they help us remember what various artifacts relate to...for example, I developed an event mapping file for Windows Event Log records, so as a tool parses through the information available, it can assign a category to various event records. This way, you no longer have to search Google or look up on a separate sheet of paper what that event refers to...you have "Login" or "Failed Login Attempt" right there next to the event description. This is particularly useful, as of Vista, Microsoft began employing a new Windows Event Log model, which means that there are a LOT more Event Logs than just the three main ones we're used to seeing. Sometimes, you'll see one event in the System or Security Event Log that will have corresponding events in other event logs, or there will be one event all by itself...knowing what these events refer to, and having a category listed for each, is extremely valuable, and I've found it to really help me a great deal with my analysis.
One way to make use of event categories is to employ an analysis matrix. What is an "analysis matrix"? Well, what happens many times is that analysts will get some general (re: "vague") analysis goals, and perhaps not really know where to start. By categorizing the various artifacts on a Windows system, we can create an analysis matrix that provides us with a means for at least begin our analysis.
An analysis matrix might appear as follows:
Again, this is simply a notional matrix, and is meant solely as an example. However, it's also a valid matrix, and something that I've used. Consider "data exfiltration"...the various categories we use to describe a "data exfiltration" case may often depend upon what you learn from a "customer" or other source. For example, I did not put an "X" in the row for "Network Access", as I have had cases where access to USB devices was specified by the customer...they felt confident that with how their infrastructure was designed that this was not an option that they wanted me to pursue. However, you may want to add this one...I have also conducted examinations in which part of what I was asked to determine was network access, such as a user taking their work laptop home and connecting to other wireless networks.
The analysis matrix is not intended to be the "be-all-end-all" of analysis, nor is it intended to be written in stone. Rather, it's intended to be something of a living document, something that provides analysts with a means for identifying what they (intend to) do, as well as serve as a foundation on which further analysis can be built. By using an analysis matrix, we have case documentation available to us immediately. An analysis matrix can also provide us with pivot points for our timeline analysis; rather than combing through thousands of records in a timeline, we now not only have a means of going after that information which may be most important to our examination, but it also helps us avoid those annoying rabbit holes that we find ourselves going down sometimes.
Finally, consider this...trying to keep track of all of the possible artifacts on a Windows system can be a daunting task. However, it can be much easier if we were to compartmentalize various artifacts into categories, making it an easier task to manage by breaking it down into smaller, easier-to-manage pieces. Rather than getting swept up in the issues surrounding a new artifact (Jump Lists are new as of Windows 7, for example...) we can simply place that artifact in the appropriate category, and incorporate it directly into our analysis.
I've talked before in the blog about how to categorize various artifacts...in fact, in this post, I talked about the different ways that Windows shortcut files can be categorized. We can look at access to USB devices as storage access, and include sub-categories for various other artifacts.
Interested in Windows DFIR training? Check it out...Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar.
For one, if we categorize events, we can raise our level of awareness of the context of the data that we're analyzing. Having categories for various artifacts can help us increase our relative level of confidence in the data that we're analyzing, because instead of looking at just one artifact, we're going to be looking at various similar, related artifacts together.
Another benefit of artifact categories is that they help us remember what various artifacts relate to...for example, I developed an event mapping file for Windows Event Log records, so as a tool parses through the information available, it can assign a category to various event records. This way, you no longer have to search Google or look up on a separate sheet of paper what that event refers to...you have "Login" or "Failed Login Attempt" right there next to the event description. This is particularly useful, as of Vista, Microsoft began employing a new Windows Event Log model, which means that there are a LOT more Event Logs than just the three main ones we're used to seeing. Sometimes, you'll see one event in the System or Security Event Log that will have corresponding events in other event logs, or there will be one event all by itself...knowing what these events refer to, and having a category listed for each, is extremely valuable, and I've found it to really help me a great deal with my analysis.
One way to make use of event categories is to employ an analysis matrix. What is an "analysis matrix"? Well, what happens many times is that analysts will get some general (re: "vague") analysis goals, and perhaps not really know where to start. By categorizing the various artifacts on a Windows system, we can create an analysis matrix that provides us with a means for at least begin our analysis.
An analysis matrix might appear as follows:
| Malware Detection | Data Exfil | Illicit Images | IP Theft | |
|---|---|---|---|---|
| Malware | X | X | ||
| Program Execution | X | X | X | |
| File Access | X | X | X | |
| Storage Access | X | X | X | |
| Network Access | X |
Again, this is simply a notional matrix, and is meant solely as an example. However, it's also a valid matrix, and something that I've used. Consider "data exfiltration"...the various categories we use to describe a "data exfiltration" case may often depend upon what you learn from a "customer" or other source. For example, I did not put an "X" in the row for "Network Access", as I have had cases where access to USB devices was specified by the customer...they felt confident that with how their infrastructure was designed that this was not an option that they wanted me to pursue. However, you may want to add this one...I have also conducted examinations in which part of what I was asked to determine was network access, such as a user taking their work laptop home and connecting to other wireless networks.
The analysis matrix is not intended to be the "be-all-end-all" of analysis, nor is it intended to be written in stone. Rather, it's intended to be something of a living document, something that provides analysts with a means for identifying what they (intend to) do, as well as serve as a foundation on which further analysis can be built. By using an analysis matrix, we have case documentation available to us immediately. An analysis matrix can also provide us with pivot points for our timeline analysis; rather than combing through thousands of records in a timeline, we now not only have a means of going after that information which may be most important to our examination, but it also helps us avoid those annoying rabbit holes that we find ourselves going down sometimes.
Finally, consider this...trying to keep track of all of the possible artifacts on a Windows system can be a daunting task. However, it can be much easier if we were to compartmentalize various artifacts into categories, making it an easier task to manage by breaking it down into smaller, easier-to-manage pieces. Rather than getting swept up in the issues surrounding a new artifact (Jump Lists are new as of Windows 7, for example...) we can simply place that artifact in the appropriate category, and incorporate it directly into our analysis.
I've talked before in the blog about how to categorize various artifacts...in fact, in this post, I talked about the different ways that Windows shortcut files can be categorized. We can look at access to USB devices as storage access, and include sub-categories for various other artifacts.
Interested in Windows DFIR training? Check it out...Timeline Analysis, 4-5 Feb; Windows Forensic Analysis, 11-12 Mar.
Tuesday, January 08, 2013
Training
For those readers who may not be aware, I teach a couple of training courses through my employer, at our facility in Reston, VA. We're also available to deliver those courses at your location, if requested. As such, I thought it might be helpful to provide some information about the courses, so in this post, I'll talk about the courses we offer, some we're looking to offer, and what you can expect to get out of the courses.
Windows Forensic Analysis
Day 1 starts with a course introduction, and then we get right into discussing some core analysis concepts, which will be addressed again and again throughout the training. From there, we begin exploring and discussing some of the various data sources and artifacts available on Windows 7 systems. Knowing that XP is still out there, we don't ignore that version of Windows, we simply focus primarily on Windows 7. Artifacts specific to other systems are discussed, as they come up.
Throughout the course, we also discuss the various artifact categories, and how to create and use an analysis matrix to focus and document your analysis. We discuss what data is available, how to get it, how to correlate that data with other available data, and how to get previous versions of that data by accessing Volume Shadow Copies. All of this is accompanied by hands-on demonstrations of tools and techniques; many of the tools used are only available to those attending the training.
Day 2 starts with a quick review of the previous day's materials and answering any questions attendees may have; if there's any material that needs to be completed from the first day, we finish up with that, and then move into the hands-on exercises. Depending upon the attendee's familiarity with the tools and techniques used, these exercises may be guided, or they will be completed by attendees, in teams or individually.
Do you want to know what secrets lie hidden within Windows shortcut files and Jump Lists? Want to know more about "shellbags"? How about other artifacts? This course will tell...no, show...you. Not only that, we'll show you how to use this information to a greater effect, in a more timely and efficient manner, in order to extend your analysis.
Each attendee receives a copy of Windows Forensic Analysis Toolkit 3/e.
Timeline Analysis
Day 1 - Much like the Windows Forensic Analysis course, we start the first day with some core analysis concepts specific to timeline analysis, and then we jump right into exploring and discussing various data sources and artifacts as they relate to creating and analyzing timelines. We discuss the various artifact and event categories, and how this information can be used to get more out of your timeline analysis.
Day 2 starts off with completing any material from the first day, answering any questions the attendees may have, and then kicking off into a series of scenarios where questions are answered based on findings from a timeline; we not only go over how to create a timeline, but also how to go about analyzing that timeline and finding the answers to the questions.
If you can't remember all of the commands that we go over in the course, don't worry...you can write down notes on the provided copies of the slides, or you can turn to the provided cheat sheet for hints and reminders. Many of the tools used in this course are only available to those attending the course.
Each attendee receives a copy of Windows Forensic Analysis Toolkit 3/e.
Registry Analysis
This 1-day course is based on the material in my book, Windows Registry Forensics. As such, we spend some time in this course discussing not only the structure of the Registry, but also the value of performing Registry analysis. There is a good deal of information in the Registry that can significantly impact your analysis, and the goal of this course is to allow you to go beyond assumption to determining explicitly why you're seeing what you're seeing.
As you would guess, we spend some time discussing various tools, and some attention is given to RegRipper. For those interested, attendees will receive plugins that are not available through the public distribution. We also spend some time discussing the RegRipper components and structure, how it's used, and how to get the most out of it.
One of the take-aways we provide with this course is a graphic illustrating various components of USB device analysis, showing artifacts that aren't addressed anywhere else.
Each attendee receives a copy of Windows Registry Forensics.
Why Should I Attend?
That's always a great question; it's one I ask myself, as well, whenever I have an option to attend training.
Each attendee is provided the tools for the course, which includes tools that are only available to you if you attend the course. Tools for parsing various data structures, including RegRipper plugins that you can't get any place else. Several publicly available tools are discussed in the courses, but due to licenses, are not provided with the course materials. In such cases, the materials provide links to the tools.
I continually update the course materials. I sit down with the materials immediately following a course and look at my notes, any questions asked by attendees, and I pay particular attention to the course evaluation forms. When something new pops up in the media, I like to be sure to include it in the course for discussion. Updates come from other areas, as well...most notably, what I get from and how I perform my analysis. New techniques and findings are continually incorporated directly into the training materials.
As the Windows operating systems have gotten more complex, it's proven to be difficult for a lot of analysts to maintain current knowledge of the various artifacts, as well as analysis tools and techniques. These courses will not only provide you with the information, but also provide you with an opportunity to use those tools and employ those techniques, developing an understanding of each so that you can incorporate them into your analysis processes.
What Do I Need To Know Before Attending?
For the currently available courses, we ask that you arrive with a laptop with Windows 7 installed (can be a VM), a familiarity with operating at the command prompt, and a desire to learn. Bring your questions. While sample data is provided with the course materials, feel free to bring your own data, if you like.
The courses are developed so that you do NOT want to book all of these courses in a single 5-day training course. The reason is that a great deal of information is provided in the Windows Forensic Analysis course, and if you've never done timeline analysis before (and in some cases, even if you have), you do not want to immediately step off into the Timeline Analysis course. It is best to take the Windows Forensic Analysis (and perhaps the Registry Analysis) course(s), return to your shop, and make develop your familiarity with the data sources before taking the Timeline Analysis course.
If you've ever seen or heard me present, you know that I am less about lecturing and more about interacting. If you're interested in engaging and interacting with others to better understand data sources and artifacts, as well as how they can be used to further your analysis, then sign up for one of our courses.
Upcoming Course(s)
Malware Detection - By request, I'm working a course that addresses malware detection within an acquired image. I've taught courses similar to this before, and I think that in a lot of ways, it's an eye-opener for a lot of folks, even those who deal with malware regularly. This is NOT a malware analysis course...the purpose of this course is to help analysts understand how to locate malware within an acquired image. This is one of those analysis skills that traverses a number of cases, from breaches to data theft, even to claims of the "Trojan Defense".
Others - TBD.
Our website includes information regarding the schedule of courses, as well as the cost for each course. Check back regularly, as the schedule may change. Also, if you're interested in having us come to you to provide the training, let us know.
Windows Forensic Analysis
Day 1 starts with a course introduction, and then we get right into discussing some core analysis concepts, which will be addressed again and again throughout the training. From there, we begin exploring and discussing some of the various data sources and artifacts available on Windows 7 systems. Knowing that XP is still out there, we don't ignore that version of Windows, we simply focus primarily on Windows 7. Artifacts specific to other systems are discussed, as they come up.
Throughout the course, we also discuss the various artifact categories, and how to create and use an analysis matrix to focus and document your analysis. We discuss what data is available, how to get it, how to correlate that data with other available data, and how to get previous versions of that data by accessing Volume Shadow Copies. All of this is accompanied by hands-on demonstrations of tools and techniques; many of the tools used are only available to those attending the training.
Day 2 starts with a quick review of the previous day's materials and answering any questions attendees may have; if there's any material that needs to be completed from the first day, we finish up with that, and then move into the hands-on exercises. Depending upon the attendee's familiarity with the tools and techniques used, these exercises may be guided, or they will be completed by attendees, in teams or individually.
Do you want to know what secrets lie hidden within Windows shortcut files and Jump Lists? Want to know more about "shellbags"? How about other artifacts? This course will tell...no, show...you. Not only that, we'll show you how to use this information to a greater effect, in a more timely and efficient manner, in order to extend your analysis.
Each attendee receives a copy of Windows Forensic Analysis Toolkit 3/e.
Timeline Analysis
Day 1 - Much like the Windows Forensic Analysis course, we start the first day with some core analysis concepts specific to timeline analysis, and then we jump right into exploring and discussing various data sources and artifacts as they relate to creating and analyzing timelines. We discuss the various artifact and event categories, and how this information can be used to get more out of your timeline analysis.
Day 2 starts off with completing any material from the first day, answering any questions the attendees may have, and then kicking off into a series of scenarios where questions are answered based on findings from a timeline; we not only go over how to create a timeline, but also how to go about analyzing that timeline and finding the answers to the questions.
If you can't remember all of the commands that we go over in the course, don't worry...you can write down notes on the provided copies of the slides, or you can turn to the provided cheat sheet for hints and reminders. Many of the tools used in this course are only available to those attending the course.
Each attendee receives a copy of Windows Forensic Analysis Toolkit 3/e.
Registry Analysis
This 1-day course is based on the material in my book, Windows Registry Forensics. As such, we spend some time in this course discussing not only the structure of the Registry, but also the value of performing Registry analysis. There is a good deal of information in the Registry that can significantly impact your analysis, and the goal of this course is to allow you to go beyond assumption to determining explicitly why you're seeing what you're seeing.
As you would guess, we spend some time discussing various tools, and some attention is given to RegRipper. For those interested, attendees will receive plugins that are not available through the public distribution. We also spend some time discussing the RegRipper components and structure, how it's used, and how to get the most out of it.
One of the take-aways we provide with this course is a graphic illustrating various components of USB device analysis, showing artifacts that aren't addressed anywhere else.
Each attendee receives a copy of Windows Registry Forensics.
Why Should I Attend?
That's always a great question; it's one I ask myself, as well, whenever I have an option to attend training.
Each attendee is provided the tools for the course, which includes tools that are only available to you if you attend the course. Tools for parsing various data structures, including RegRipper plugins that you can't get any place else. Several publicly available tools are discussed in the courses, but due to licenses, are not provided with the course materials. In such cases, the materials provide links to the tools.
I continually update the course materials. I sit down with the materials immediately following a course and look at my notes, any questions asked by attendees, and I pay particular attention to the course evaluation forms. When something new pops up in the media, I like to be sure to include it in the course for discussion. Updates come from other areas, as well...most notably, what I get from and how I perform my analysis. New techniques and findings are continually incorporated directly into the training materials.
As the Windows operating systems have gotten more complex, it's proven to be difficult for a lot of analysts to maintain current knowledge of the various artifacts, as well as analysis tools and techniques. These courses will not only provide you with the information, but also provide you with an opportunity to use those tools and employ those techniques, developing an understanding of each so that you can incorporate them into your analysis processes.
What Do I Need To Know Before Attending?
For the currently available courses, we ask that you arrive with a laptop with Windows 7 installed (can be a VM), a familiarity with operating at the command prompt, and a desire to learn. Bring your questions. While sample data is provided with the course materials, feel free to bring your own data, if you like.
The courses are developed so that you do NOT want to book all of these courses in a single 5-day training course. The reason is that a great deal of information is provided in the Windows Forensic Analysis course, and if you've never done timeline analysis before (and in some cases, even if you have), you do not want to immediately step off into the Timeline Analysis course. It is best to take the Windows Forensic Analysis (and perhaps the Registry Analysis) course(s), return to your shop, and make develop your familiarity with the data sources before taking the Timeline Analysis course.
If you've ever seen or heard me present, you know that I am less about lecturing and more about interacting. If you're interested in engaging and interacting with others to better understand data sources and artifacts, as well as how they can be used to further your analysis, then sign up for one of our courses.
Upcoming Course(s)
Malware Detection - By request, I'm working a course that addresses malware detection within an acquired image. I've taught courses similar to this before, and I think that in a lot of ways, it's an eye-opener for a lot of folks, even those who deal with malware regularly. This is NOT a malware analysis course...the purpose of this course is to help analysts understand how to locate malware within an acquired image. This is one of those analysis skills that traverses a number of cases, from breaches to data theft, even to claims of the "Trojan Defense".
Others - TBD.
Our website includes information regarding the schedule of courses, as well as the cost for each course. Check back regularly, as the schedule may change. Also, if you're interested in having us come to you to provide the training, let us know.
Saturday, January 05, 2013
There Are Four Lights: USB-Accessible Storage
There's been a good deal of discussion and documentation regarding discovering USB devices that had been connected to a Windows system, as this seems to be very important to a number of examiners. In 2005, Cory Altheide and I published some initial information, and over the years since then, that information has been expanded, simply because it continues to grow. For example, Rob Lee has published valuable checklists via the SANS Forensics Blog, and Jacky Fox recently published her dissertation, which includes some interesting and valuable information regarding interpreting some of the information that is available regarding user access to USB devices via the Registry. Ms. Fox determined that when a USB device is connected to a system and mounted as a volume, that volume GUID is added to the MountPoints2 key for all logged in users, not just the user logged in at the console.
Further, Mark Woan recently updated information collected by his USBDeviceForensics tool, to include querying some additional keys/values.
Regarding the additional keys/values that Mark's tool is querying, Windows 7 and 8 systems have additional values beneath the device keys in the System hive, specifically a "Property" key with a number of GUID subkeys. This blog post provides some very good information that facilitates further searches, which leads use to information regarding a time stamp value that pertains to the InstallDate, as well as one that pertains to the FirstInstallDate.
So what? Well, let's take a look at the MS definition for the FirstInstallDate:
Windows sets the value of DEVPKEY_Device_FirstInstallDate with the time stamp that specifies when the device instance was first installed in the system.
Pretty cool, eh? This is what MS says about the InstallDate time stamp:
This time stamp value changes for each successive update of the device driver. For example, this time stamp reports the date and time when the device driver was last updated through Windows Update.
Ah, interesting. So it would appear that, based on the MS definitions for these values, we now have the information about when the device was first connected to the system available right there in the Registry. I'm not saying that we don't have to go anywhere else...rather, I'm suggesting that we have corroborating data that we can use to provide an increased relative confidence (a phrase that you usually see in my posts regarding timelines) in the data that we're analyzing.
Something that hasn't been addressed is that most of the publicly-available processes that are currently being used are not as complete as they could be. Wait...what? Well, this is where specificity of language within the DFIR community comes into play...it turns out that the processes are actually really very good, as long as all we're interested in is specifically USB thumb drives or external drives. However, there are devices that can be connected to Windows systems via USB and accessed as storage devices (digital cameras, iStuff, smartphone handsets), that do not necessarily become apparent to analysts using the commonly-accepted tools, processes and checklists. We can find these devices by looking beneath other Registry keys, as well as in other locations beyond the Registry, and by correlating information between them. This is particularly useful when counter-forensics techniques have been used (however unintentional...), as not everything may be completely gone, and we may be able to find some remnant (LNK file, shellbags, deleted Registry keys/values, Windows Event Log, etc.) that will point us to the use of such devices.
One of the pitfalls of interpretation of Registry data, as Ms. Fox pointed out in her dissertation, is that we often don't have current, up-to-date databases of all devices that could be connected to a Windows system, so we might see vendor ID (VID) and product ID (PID) values within key names beneath the Enum\USB key, but not know what they translate to...I've found Motorola devices, for instance, that required a good deal of searching in order to determine which smartphone handset was pointed to by the PID value. As such, no process is going to be 100%, push-a-button complete, but the point is that we will know that the data is there, we know to get it, and we know how to use it.
Full analysis of USB-accessible storage media can be extremely important to a number of exams, such as illicit image and IP theft cases. Many examiners used to think that sneaking a thumb drive into an infrastructure was a threat...and it still is; these devices get smaller and smaller every day, while their capacity increases. But we need to start thinking about other USB-accessible storage, such as smartphones and iDevices, not because they're easily hidden, but because they're so ubiquitous that we tend to not focus on them...we take them for granted.
A Mapping Technique
The EMDMgmt subkey (within the Software Registry hive) names include the serial number for the mounted volume (VSN), which is also included in the MS-SHLLLINK structure, which itself is found in Windows shortcut/LNK files, as well as Windows 7 and 8 Jump Lists. By correlating the VSNs from multiple sources, I was able to illustrate access to external storage devices in a manner that overcomes the shortcoming identified by Ms. Fox. What I've done is used code to parse through the LNK structures (LNK files in the Recent folder, for example, and the LNK streams within the Jump Lists) to list the VSNs, looking for the one (or two, or however many...) that point to the device identified in the EMDMgmt subkey name.
Further, Mark Woan recently updated information collected by his USBDeviceForensics tool, to include querying some additional keys/values.
Regarding the additional keys/values that Mark's tool is querying, Windows 7 and 8 systems have additional values beneath the device keys in the System hive, specifically a "Property" key with a number of GUID subkeys. This blog post provides some very good information that facilitates further searches, which leads use to information regarding a time stamp value that pertains to the InstallDate, as well as one that pertains to the FirstInstallDate.
So what? Well, let's take a look at the MS definition for the FirstInstallDate:
Windows sets the value of DEVPKEY_Device_FirstInstallDate with the time stamp that specifies when the device instance was first installed in the system.
Pretty cool, eh? This is what MS says about the InstallDate time stamp:
This time stamp value changes for each successive update of the device driver. For example, this time stamp reports the date and time when the device driver was last updated through Windows Update.
Ah, interesting. So it would appear that, based on the MS definitions for these values, we now have the information about when the device was first connected to the system available right there in the Registry. I'm not saying that we don't have to go anywhere else...rather, I'm suggesting that we have corroborating data that we can use to provide an increased relative confidence (a phrase that you usually see in my posts regarding timelines) in the data that we're analyzing.
Something that hasn't been addressed is that most of the publicly-available processes that are currently being used are not as complete as they could be. Wait...what? Well, this is where specificity of language within the DFIR community comes into play...it turns out that the processes are actually really very good, as long as all we're interested in is specifically USB thumb drives or external drives. However, there are devices that can be connected to Windows systems via USB and accessed as storage devices (digital cameras, iStuff, smartphone handsets), that do not necessarily become apparent to analysts using the commonly-accepted tools, processes and checklists. We can find these devices by looking beneath other Registry keys, as well as in other locations beyond the Registry, and by correlating information between them. This is particularly useful when counter-forensics techniques have been used (however unintentional...), as not everything may be completely gone, and we may be able to find some remnant (LNK file, shellbags, deleted Registry keys/values, Windows Event Log, etc.) that will point us to the use of such devices.
One of the pitfalls of interpretation of Registry data, as Ms. Fox pointed out in her dissertation, is that we often don't have current, up-to-date databases of all devices that could be connected to a Windows system, so we might see vendor ID (VID) and product ID (PID) values within key names beneath the Enum\USB key, but not know what they translate to...I've found Motorola devices, for instance, that required a good deal of searching in order to determine which smartphone handset was pointed to by the PID value. As such, no process is going to be 100%, push-a-button complete, but the point is that we will know that the data is there, we know to get it, and we know how to use it.
Full analysis of USB-accessible storage media can be extremely important to a number of exams, such as illicit image and IP theft cases. Many examiners used to think that sneaking a thumb drive into an infrastructure was a threat...and it still is; these devices get smaller and smaller every day, while their capacity increases. But we need to start thinking about other USB-accessible storage, such as smartphones and iDevices, not because they're easily hidden, but because they're so ubiquitous that we tend to not focus on them...we take them for granted.
A Mapping Technique
The EMDMgmt subkey (within the Software Registry hive) names include the serial number for the mounted volume (VSN), which is also included in the MS-SHLLLINK structure, which itself is found in Windows shortcut/LNK files, as well as Windows 7 and 8 Jump Lists. By correlating the VSNs from multiple sources, I was able to illustrate access to external storage devices in a manner that overcomes the shortcoming identified by Ms. Fox. What I've done is used code to parse through the LNK structures (LNK files in the Recent folder, for example, and the LNK streams within the Jump Lists) to list the VSNs, looking for the one (or two, or however many...) that point to the device identified in the EMDMgmt subkey name.
Tuesday, January 01, 2013
BinMode
I've recently been working on a script to parse the NTFS $UsnJrnl:$J file, also known as the USN Change Journal. Rather than blogging about the technical aspects of what this file is, or why a forensic analyst would want to parse it, I thought that this would be a great opportunity to instead talk about programming and parsing binary structures.
There are several things I like about being able to program, as an aspect of my DFIR work:
- It very often allows me to achieve something that I cannot achieve through the use of commercially available tools. Sometimes it allows me to "get there" faster, other times, it's the only way to "get there".
- I tend to break my work down into distinct, compartmentalized tasks, which lends itself well to programming (and vice versa).
- It gives me a challenge. I can focus my effort and concentration on solving a problem, one that I will likely see again and will already have an automated solution for solving when I see it.
- It allows me to see the data in its raw form, not filtered through an application written by a developer. This allows me to see data within the various structures (based on structure definitions from MS and others), and possibly find new ways to use that data.
One of the benefits of programming is that I have all of this code available, not just as complete applications but also stuff I've written to help me perform analysis. Stuff like translating time values (FILETIME objects, DOSDate time stamps, etc.), as well as a printData() function that takes binary data of an arbitrary length and translates it into a hex editor-style view, which makes it easy to print out sections of data and work with them directly. Being able to reuse this code (even if "code reuse" is simply a matter of copy-paste) means that I can achieve a pretty extensive depth of analysis in fairly short order, reducing the time it takes for me to collect, parse, and analyze data at a more comprehensive level than before. If I'm parsing some data, and use the printData() function to display the binary data in hex at the console, I may very well recognize a 64-bit time stamp at a regular offset, and then be able to add that to my parsing routine. That's kind of how I went about writing the shellbags.pl plugin for RegRipper.
I've also recently been looking at IE index.dat files in a hex editor, and writing my own parser based on the MSIE Cache File Format put together by Joachim Metz. So far, my initial parser works very well against the index.dat file in the TIF folder, as well as the one associated with the cookies. But what's really fascinating about this is what I'm seeing...each record has two FILETIME objects and up to three DOSDate (aka, FATTime) time stamps, in addition to other metadata. For any given entry, all of these fields may not be populated, but the fact is that I can view them...and verify them with a hex editor, if necessary.
As a side note regarding that code, I've found it very useful so far. I can run the code at the command line, and pipe the output through one or more "find" commands in order to locate or view specific entries. For example, the following command line gets the "Location : " fields for me, and then looks for specific entries; in this case, "apple":
C:\tools>parseie.pl index.dat | find "Location :" | find "apple" /i
Using the above command line, I'm able to narrow down the access to specific things, such as purchase of items via the Apple Store, etc.
I've also been working on a $UsnJrnl (actually, the $UsnJrnl:$J ADS file) parser, which itself has been fascinating. This work was partially based on something I've felt that I've needed to do for a while now, and talking to Corey Harrell about some of his recent findings has renewed my interest in this effort, particularly as it applies to malware detection.
Understanding binary structures can be very helpful. For example, consider the target.lnk file illustrated in this write-up of the Gauss malware. If you parse the information manually, using the MS specification...which should not be hard because there are only 0xC3 bytes visible...you'll see that the FILETIME time stamps for the target file are nonsense (Cheeky4n6Monkey got that, as well). As you parse the shell item ID list, based on the MS specification, you'll see that the first item is a System folder that points to "My Computer", and the second item is a Device entry whose GUID is "{21ec2020-3aea-1069-a2dd-08002b30309d}". When I looked this GUID up online, I found some interesting references to protecting or locking folders, such as this one at LIUtilities, and this one at GovernmentSecurity.org. I found this list of shell folder IDs, which might also be useful.
The final shell item, located at offset 0x84, is type 0x06, which isn't something that I've seen before. But there's nothing in the write-up that explains in detail how this LNK file might be used by the malware for persistence or propagation, so this was just an interesting exercise for me, as well as for Cheeky4n6Monkey, who also worked on parsing the target.lnk file manually. So, why even bother? Well, like I said, it's extremely beneficial to understand the format of various binary structures, but there's another reason. Have you read these posts over on the CyanLab blog? No? You should. I've seen shortcut/LNK files with no LinkInfo block, only the shell item ID list, that point to devices; as such, being able to parse and understand these...or even just recognize them...can be very beneficial if you're at all interested in determining USB storage devices that had been connected to a system. So far, most of these devices that I have seen have been digital cameras and smart phone handsets.
Everything
Okay, right about now, you're probably thinking, "so what?" Who cares, right? Well, this should be a very interesting, if not outright important issue for DFIR analysts...many of whom want to see everything when it comes to analysis. So the question then becomes...are you seeing everything? When you run your tool of choice, is it getting everything?
Folks like Chris Pogue talk a lot about analysis techniques like "sniper forensics", which is an extremely valuable means for performing data collection and analysis. However, let's take another look at the above question, from the perspective of sniper forensics...do you have the data you need? If you don't know what's there, how do you know?
If you don't know that Windows shortcut files include a shell item ID list, and what that data means, then how can you evaluate the use of a tool that parses LNK files? I'm using shell item ID lists as an example, simply because they're so very pervasive on Windows 7 systems...they're in shortcut files, Jump Lists, Registry value data. They're in a LOT of Registry value data. But the concept applies to other aspects of analysis, such as browser analysis. When you're performing browser analysis in order to determine user activity, are you just checking the history and cookies, or are you including Registry settings ("TypedURLs" key values for IE 5-9, and "TypedURLsTimes" key values on Windows 8), bookmarks, and session restore files? When performing USB device analysis on Windows systems, are you looking for all devices, or are you using checklists that only cover thumb drives and external hard drives?
I know that my previous paragraph covers a couple of different levels of granularity, but the point remains the same...are you getting everything that you need or want to perform your analysis? Does the tool you're using get all system and/or user activity, or does it get some of it?
Can we ever know it all?
One of the aspects of the DFIR community is that, for the most part, most of us seem to work in isolation. We work our cases and exams, and don't really bother too much with asking someone else, someone we know and trust, "hey, did I look at everything I could have here?" or "did I look at everything I needed to in order to address my analysis goals in a comprehensive manner?" For a variety of reasons, we don't tend to seek out peer review, even after cases are over and done.
But you know something...we can't know it all. No one of us is as smart or experienced as several or all of us working together. This can be close collaboration, face-to-face, or online collaboration through blogs, or sites such as the ForensicsWiki, which makes a great repository, if it's used.
Choices
Finally, a word about choices in programming languages to use. Some folks have a preference. I've been using Perl for a long time, since 1999. I learned BASIC in the '80s, as well as some Pascal, and then in the mid-'90s, I picked up some Java as part of my graduate studies. I know some folks prefer Python, and that's fine. Some folks within the community would like to believe that there are sharp divides between these two camps, that some who use one language detest the other, as well as those who use it. Nothing could be further from the truth. In fact, I would suggest that this attempt to create drama where there is none is simply a means of masking the fact that some analysts and examiners simply don't understand the technical aspects of the work that's actually being done.
Resources
Forensics from the Sausage Factory - USN Change Journal
Security BrainDump - Post regarding the USN Change Journal
OpenFoundry - Free tools; page includes link to a Python script for parsing the $UsnJrnl:$J file
There are several things I like about being able to program, as an aspect of my DFIR work:
- It very often allows me to achieve something that I cannot achieve through the use of commercially available tools. Sometimes it allows me to "get there" faster, other times, it's the only way to "get there".
- I tend to break my work down into distinct, compartmentalized tasks, which lends itself well to programming (and vice versa).
- It gives me a challenge. I can focus my effort and concentration on solving a problem, one that I will likely see again and will already have an automated solution for solving when I see it.
- It allows me to see the data in its raw form, not filtered through an application written by a developer. This allows me to see data within the various structures (based on structure definitions from MS and others), and possibly find new ways to use that data.
One of the benefits of programming is that I have all of this code available, not just as complete applications but also stuff I've written to help me perform analysis. Stuff like translating time values (FILETIME objects, DOSDate time stamps, etc.), as well as a printData() function that takes binary data of an arbitrary length and translates it into a hex editor-style view, which makes it easy to print out sections of data and work with them directly. Being able to reuse this code (even if "code reuse" is simply a matter of copy-paste) means that I can achieve a pretty extensive depth of analysis in fairly short order, reducing the time it takes for me to collect, parse, and analyze data at a more comprehensive level than before. If I'm parsing some data, and use the printData() function to display the binary data in hex at the console, I may very well recognize a 64-bit time stamp at a regular offset, and then be able to add that to my parsing routine. That's kind of how I went about writing the shellbags.pl plugin for RegRipper.
I've also recently been looking at IE index.dat files in a hex editor, and writing my own parser based on the MSIE Cache File Format put together by Joachim Metz. So far, my initial parser works very well against the index.dat file in the TIF folder, as well as the one associated with the cookies. But what's really fascinating about this is what I'm seeing...each record has two FILETIME objects and up to three DOSDate (aka, FATTime) time stamps, in addition to other metadata. For any given entry, all of these fields may not be populated, but the fact is that I can view them...and verify them with a hex editor, if necessary.
As a side note regarding that code, I've found it very useful so far. I can run the code at the command line, and pipe the output through one or more "find" commands in order to locate or view specific entries. For example, the following command line gets the "Location : " fields for me, and then looks for specific entries; in this case, "apple":
C:\tools>parseie.pl index.dat | find "Location :" | find "apple" /i
Using the above command line, I'm able to narrow down the access to specific things, such as purchase of items via the Apple Store, etc.
I've also been working on a $UsnJrnl (actually, the $UsnJrnl:$J ADS file) parser, which itself has been fascinating. This work was partially based on something I've felt that I've needed to do for a while now, and talking to Corey Harrell about some of his recent findings has renewed my interest in this effort, particularly as it applies to malware detection.
Understanding binary structures can be very helpful. For example, consider the target.lnk file illustrated in this write-up of the Gauss malware. If you parse the information manually, using the MS specification...which should not be hard because there are only 0xC3 bytes visible...you'll see that the FILETIME time stamps for the target file are nonsense (Cheeky4n6Monkey got that, as well). As you parse the shell item ID list, based on the MS specification, you'll see that the first item is a System folder that points to "My Computer", and the second item is a Device entry whose GUID is "{21ec2020-3aea-1069-a2dd-08002b30309d}". When I looked this GUID up online, I found some interesting references to protecting or locking folders, such as this one at LIUtilities, and this one at GovernmentSecurity.org. I found this list of shell folder IDs, which might also be useful.
The final shell item, located at offset 0x84, is type 0x06, which isn't something that I've seen before. But there's nothing in the write-up that explains in detail how this LNK file might be used by the malware for persistence or propagation, so this was just an interesting exercise for me, as well as for Cheeky4n6Monkey, who also worked on parsing the target.lnk file manually. So, why even bother? Well, like I said, it's extremely beneficial to understand the format of various binary structures, but there's another reason. Have you read these posts over on the CyanLab blog? No? You should. I've seen shortcut/LNK files with no LinkInfo block, only the shell item ID list, that point to devices; as such, being able to parse and understand these...or even just recognize them...can be very beneficial if you're at all interested in determining USB storage devices that had been connected to a system. So far, most of these devices that I have seen have been digital cameras and smart phone handsets.
Everything
Okay, right about now, you're probably thinking, "so what?" Who cares, right? Well, this should be a very interesting, if not outright important issue for DFIR analysts...many of whom want to see everything when it comes to analysis. So the question then becomes...are you seeing everything? When you run your tool of choice, is it getting everything?
Folks like Chris Pogue talk a lot about analysis techniques like "sniper forensics", which is an extremely valuable means for performing data collection and analysis. However, let's take another look at the above question, from the perspective of sniper forensics...do you have the data you need? If you don't know what's there, how do you know?
If you don't know that Windows shortcut files include a shell item ID list, and what that data means, then how can you evaluate the use of a tool that parses LNK files? I'm using shell item ID lists as an example, simply because they're so very pervasive on Windows 7 systems...they're in shortcut files, Jump Lists, Registry value data. They're in a LOT of Registry value data. But the concept applies to other aspects of analysis, such as browser analysis. When you're performing browser analysis in order to determine user activity, are you just checking the history and cookies, or are you including Registry settings ("TypedURLs" key values for IE 5-9, and "TypedURLsTimes" key values on Windows 8), bookmarks, and session restore files? When performing USB device analysis on Windows systems, are you looking for all devices, or are you using checklists that only cover thumb drives and external hard drives?
I know that my previous paragraph covers a couple of different levels of granularity, but the point remains the same...are you getting everything that you need or want to perform your analysis? Does the tool you're using get all system and/or user activity, or does it get some of it?
Can we ever know it all?
One of the aspects of the DFIR community is that, for the most part, most of us seem to work in isolation. We work our cases and exams, and don't really bother too much with asking someone else, someone we know and trust, "hey, did I look at everything I could have here?" or "did I look at everything I needed to in order to address my analysis goals in a comprehensive manner?" For a variety of reasons, we don't tend to seek out peer review, even after cases are over and done.
But you know something...we can't know it all. No one of us is as smart or experienced as several or all of us working together. This can be close collaboration, face-to-face, or online collaboration through blogs, or sites such as the ForensicsWiki, which makes a great repository, if it's used.
Choices
Finally, a word about choices in programming languages to use. Some folks have a preference. I've been using Perl for a long time, since 1999. I learned BASIC in the '80s, as well as some Pascal, and then in the mid-'90s, I picked up some Java as part of my graduate studies. I know some folks prefer Python, and that's fine. Some folks within the community would like to believe that there are sharp divides between these two camps, that some who use one language detest the other, as well as those who use it. Nothing could be further from the truth. In fact, I would suggest that this attempt to create drama where there is none is simply a means of masking the fact that some analysts and examiners simply don't understand the technical aspects of the work that's actually being done.
Resources
Forensics from the Sausage Factory - USN Change Journal
Security BrainDump - Post regarding the USN Change Journal
OpenFoundry - Free tools; page includes link to a Python script for parsing the $UsnJrnl:$J file
Friday, December 28, 2012
Malware Detection
Corey recently posted to his blog regarding his exercise of infecting a system with ZeroAccess. In his post, Corey provides a great example of a very valuable malware artifact, as well as an investigative process, that can lead to locating malware that may be missed by more conventional means.

This post isn't meant to take anything away from Corey's exceptional work; rather, my intention is to show another perspective of the data, sort of like "The Secret PoliceMan's Other Ball". Corey's always done a fantastic job of performing research and presenting his findings, and it is not my intention to detract from his work at all. Instead, I would like to present another perspective, utilizing Corey's work and blog post as a basis, and as a stepping stone.
The ZA sample that Corey looked at was a bit different from what James Wyke of SophosLabs wrote about, but there were enough commonalities that some artifacts could be used to create an IOC or plugin for detecting the presence of this bit of malware, even if AV didn't detect it. Specifically, the file "services.exe" was infected, an EA attribute was added to the file record in the MFT, and a Registry modification occurred in order to create a persistence mechanism for the malware. Looking at these commonalities is similar to looking at the commonalities between various versions of the Conficker family, which created a randomly-named service for persistence.
From the Registry hives from Corey's test, I was able to create and test a RegRipper plugin that do a pretty good job of filtering through the Classes/CLSID subkey (from the Software hive) and locating anomalies. In it's original form, the MFT parser that I wrote finds the EA attribute, but doesn't specifically flag on it, and it can't extract the shell code and the malware PE file (because the data is non-resident). However, there were a couple of interesting things I got from parsing the MFT...
If you refer to Corey's post, take a look at the section regarding the MFT record for the infected services.exe file. If you look at the time stamps and compare those from the $STANDARD_INFORMATION attribute to those of the $FILE_NAME attribute that Corey posted, you'll see an excellent example of file system tunneling. I've talked about this in a number of my presentations, but here's a pretty cool to see an actual example of it. I know that this isn't really "outside the lab", per se, but still, it's pretty cool to this functionality as a result of a sample of malware, rather than a contrived exercise. Hopefully, this example will go a long way toward helping analysts understand what they're seeing in the time stamps.
Corey also illustrated an excellent use of timeline analysis to locate other files that were created or modified around the same time that the services.exe file was infected. What the timeline doesn't show clearly is that the time stamps were extracted from the $FILE_NAME attribute in the MFT...the $STANDARD_INFORMATION attributes for those same files indicate that there was some sort of time stamp manipulation ("timestomping") that occurred, as many of the files have M, A, and B times from 13 and 14 Jul 2009. However, the date in question that Corey looked at in his blog post was 6 Dec 2012 (the day of the test). Incorporating Prefetch file metadata and Registry key LastWrite times into a timeline would a pretty tight "grouping" of these artifacts at or "near" the same time.
Another interesting finding in analyzing the MFT is that the "new" services.exe file was MFT record number 42756 (see Corey's blog entry for the original file's record number). Looking "near" the MFT record number, there are a number of files and folders that are created (and "timestomped") prior to the new services.exe file record being created. Searching for some of the filenames and paths (such as C:\Windows\Temp\fwtsqmfile00.sqm), I find references to other variants of ZeroAccess. But what is very interesting about this is the relatively tight grouping of the file and folder creations, not based on time stamps or time stamp anomalies, but instead based on MFT record numbers.
Some take-aways from this...at least what I took away...are:
1. Timeline analysis is an extremely powerful analysis technique because it provides us with context, as well as an increased relative level of confidence in the data we're analyzing.
2. Timeline analysis can be even more powerful when it is not the sole analysis technique, but incorporated into an overall analysis plan. What about that Prefetch file for services.exe? A little bit of Prefetch file analysis would have produced some very interesting results, and using what was found through this analysis technique would have lead to other artifacts that should be examined in the timeline. Artifacts found outside of timeline analysis could be used as search terms or pivot points in a timeline, which would then provide context to those artifacts, which could then be incorporated back into other analyses.
3. Some folks have told me that having multiple tools for creating timelines makes creating timelines too complex a task; however, the tools I tend to create and use are multi-purpose. For example, I use pref.pl (I also have a 'compiled' EXE) for Prefetch file analysis, as well as parsing Prefetch file metadata into a timeline. I use RegRipper for parsing (and some modicum of analysis) of Registry hives, as well as to generate timeline data from a number of keys and value data. I find this to be extremely valuable...I can run a tool, find something interesting in a data set as a result of the analysis, and then run the tool again, against the same data set, but with a different set of switches, and populate my timeline. I don't need to switch GUIs and swap out dongles. Also, it's easy to remember the various tools and switches because (a) each tool is capable of displaying its syntax via '-h', and (b) I created a cheat sheet for the tool usage.
4. Far too often, a root cause analysis, or RCA, is not performed, for whatever reason. We're losing access to a great deal of data, and as a result, we're missing out on a great deal of intel. Intel such as, "hey, what this AV vendor wrote is good, but I tested a different sample and found this...". Perhaps the reason for not performing the RCA is that "it's too difficult", "it takes too long", or "it's not worth the effort". Well, consider my previous post, Mr. CEO...without an RCA, are you being served? What are you reporting to the board or to the SEC, and is it correct? Are you going with, "it's correct to the best of my knowledge", after you went to "Joe's Computer Forensics and Crabshack" to get the work done?
Now, to add to all of the above, take a look at this post from the Sploited blog, entitled Timeline Pivot Points with the Malware Domain List. This post provides an EXCELLENT example of how timeline analysis can be used to augment other forms of analysis, or vice versa. The post also illustrates how this sort of analysis can easily be automated. In fact, this can be part of the timeline creation mechanism....when any data source is parsed (i.e., browser history list, TypedUrls Registry key, shellbags, etc.) have any URLs extracted run in comparison to the MDL, and then generate a flag of some kind within the timeline events file, so that the flag "lives" with the event. That way, you can search for those events (based on the flag) after the timeline is created, or, as part of your analysis, create a timeline of only those events. This would be similar to scanning all files in the Temp and system32 folders, looking for PE files with odd headers or mismatched extensions, and then flagging them in the timeline, as well.
Great work to both Corey and Sploited for their posts!
This post isn't meant to take anything away from Corey's exceptional work; rather, my intention is to show another perspective of the data, sort of like "The Secret PoliceMan's Other Ball". Corey's always done a fantastic job of performing research and presenting his findings, and it is not my intention to detract from his work at all. Instead, I would like to present another perspective, utilizing Corey's work and blog post as a basis, and as a stepping stone.
The ZA sample that Corey looked at was a bit different from what James Wyke of SophosLabs wrote about, but there were enough commonalities that some artifacts could be used to create an IOC or plugin for detecting the presence of this bit of malware, even if AV didn't detect it. Specifically, the file "services.exe" was infected, an EA attribute was added to the file record in the MFT, and a Registry modification occurred in order to create a persistence mechanism for the malware. Looking at these commonalities is similar to looking at the commonalities between various versions of the Conficker family, which created a randomly-named service for persistence.
From the Registry hives from Corey's test, I was able to create and test a RegRipper plugin that do a pretty good job of filtering through the Classes/CLSID subkey (from the Software hive) and locating anomalies. In it's original form, the MFT parser that I wrote finds the EA attribute, but doesn't specifically flag on it, and it can't extract the shell code and the malware PE file (because the data is non-resident). However, there were a couple of interesting things I got from parsing the MFT...
If you refer to Corey's post, take a look at the section regarding the MFT record for the infected services.exe file. If you look at the time stamps and compare those from the $STANDARD_INFORMATION attribute to those of the $FILE_NAME attribute that Corey posted, you'll see an excellent example of file system tunneling. I've talked about this in a number of my presentations, but here's a pretty cool to see an actual example of it. I know that this isn't really "outside the lab", per se, but still, it's pretty cool to this functionality as a result of a sample of malware, rather than a contrived exercise. Hopefully, this example will go a long way toward helping analysts understand what they're seeing in the time stamps.
Corey also illustrated an excellent use of timeline analysis to locate other files that were created or modified around the same time that the services.exe file was infected. What the timeline doesn't show clearly is that the time stamps were extracted from the $FILE_NAME attribute in the MFT...the $STANDARD_INFORMATION attributes for those same files indicate that there was some sort of time stamp manipulation ("timestomping") that occurred, as many of the files have M, A, and B times from 13 and 14 Jul 2009. However, the date in question that Corey looked at in his blog post was 6 Dec 2012 (the day of the test). Incorporating Prefetch file metadata and Registry key LastWrite times into a timeline would a pretty tight "grouping" of these artifacts at or "near" the same time.
Another interesting finding in analyzing the MFT is that the "new" services.exe file was MFT record number 42756 (see Corey's blog entry for the original file's record number). Looking "near" the MFT record number, there are a number of files and folders that are created (and "timestomped") prior to the new services.exe file record being created. Searching for some of the filenames and paths (such as C:\Windows\Temp\fwtsqmfile00.sqm), I find references to other variants of ZeroAccess. But what is very interesting about this is the relatively tight grouping of the file and folder creations, not based on time stamps or time stamp anomalies, but instead based on MFT record numbers.
Some take-aways from this...at least what I took away...are:
1. Timeline analysis is an extremely powerful analysis technique because it provides us with context, as well as an increased relative level of confidence in the data we're analyzing.
2. Timeline analysis can be even more powerful when it is not the sole analysis technique, but incorporated into an overall analysis plan. What about that Prefetch file for services.exe? A little bit of Prefetch file analysis would have produced some very interesting results, and using what was found through this analysis technique would have lead to other artifacts that should be examined in the timeline. Artifacts found outside of timeline analysis could be used as search terms or pivot points in a timeline, which would then provide context to those artifacts, which could then be incorporated back into other analyses.
3. Some folks have told me that having multiple tools for creating timelines makes creating timelines too complex a task; however, the tools I tend to create and use are multi-purpose. For example, I use pref.pl (I also have a 'compiled' EXE) for Prefetch file analysis, as well as parsing Prefetch file metadata into a timeline. I use RegRipper for parsing (and some modicum of analysis) of Registry hives, as well as to generate timeline data from a number of keys and value data. I find this to be extremely valuable...I can run a tool, find something interesting in a data set as a result of the analysis, and then run the tool again, against the same data set, but with a different set of switches, and populate my timeline. I don't need to switch GUIs and swap out dongles. Also, it's easy to remember the various tools and switches because (a) each tool is capable of displaying its syntax via '-h', and (b) I created a cheat sheet for the tool usage.
4. Far too often, a root cause analysis, or RCA, is not performed, for whatever reason. We're losing access to a great deal of data, and as a result, we're missing out on a great deal of intel. Intel such as, "hey, what this AV vendor wrote is good, but I tested a different sample and found this...". Perhaps the reason for not performing the RCA is that "it's too difficult", "it takes too long", or "it's not worth the effort". Well, consider my previous post, Mr. CEO...without an RCA, are you being served? What are you reporting to the board or to the SEC, and is it correct? Are you going with, "it's correct to the best of my knowledge", after you went to "Joe's Computer Forensics and Crabshack" to get the work done?
Now, to add to all of the above, take a look at this post from the Sploited blog, entitled Timeline Pivot Points with the Malware Domain List. This post provides an EXCELLENT example of how timeline analysis can be used to augment other forms of analysis, or vice versa. The post also illustrates how this sort of analysis can easily be automated. In fact, this can be part of the timeline creation mechanism....when any data source is parsed (i.e., browser history list, TypedUrls Registry key, shellbags, etc.) have any URLs extracted run in comparison to the MDL, and then generate a flag of some kind within the timeline events file, so that the flag "lives" with the event. That way, you can search for those events (based on the flag) after the timeline is created, or, as part of your analysis, create a timeline of only those events. This would be similar to scanning all files in the Temp and system32 folders, looking for PE files with odd headers or mismatched extensions, and then flagging them in the timeline, as well.
Great work to both Corey and Sploited for their posts!
Subscribe to:
Posts (Atom)