Friday, 8 April 2011

Data Management Planning?

Guest blogger: Andrew Sallans

Following on Tom's generous invitation to write a post for the AIMS partner blog, I am finally getting around to doing so. Tom and I have been holding monthly discussions about our respective projects since sometime last summer, and have talked in great length about the commonalities between what my group (the Scientific Data Consulting Group) is dealing with in regards to research data management versus what the AIMS group is dealing with in terms of born-digital archive material.

http://www.lib.virginia.edu/


We have found that there are many areas of similarity, and that we face many of the same challenges, although we approach the problem quite differently and of course have entirely different terminology given our relative perspectives.

To get started, I have a pretty good understanding of the born-digital problem set, but have not been keeping detailed notes on the workflows and solutions that the AIMS group has identified as best practices throughout the life of this project. My intention for this post is to share the issues that we are dealing with in research data management and try to make some suggestions around areas where there may be overlap and opportunities for great information sharing and collaboration.

Starting this past January 18, 2011, the National Science Foundation (NSF) put into effect a new implementation of their pre-existing data management planning requirement. This revision now requires that researchers submit a 2 page data management plan (DMP) that specifies the steps they will take to share the data that underlies their published results. This DMP will undergo formal peer-review, require reporting in interim/final reports, and all future proposals. In effect, what one says must then be done, or else one runs the risk of losing future funding opportunities or worse, losing all funding for the institution from that particular agency. Although this requirement is focused on data sharing, it isn't possible for such an initiative to succeed without first addressing a mass of other data management issues, ranging from technical, to policy, to cultural. As we often point out these days, it is far easier to improve the process of data management up-front, in the operational process phase, than it is to begin thinking about how to share the data at the end of the project. I would expect that those attacking the born-digital archive problem can fully relate.

Here in the Scientific Data Consulting (SciDaC) Group in the UVA Library, we have been collecting and developing our local set of data management best practices for some time and have served as advisors to researchers in both the research data management and DMP development areas (they are of course interrelated, but sometimes have different levels of urgency). In doing so, we have developed what we call a "data interview/assessment" (based in large part upon the visionary work of others, Purdue's Data Curation Profile and work from the UK's Digital Curation Centre, to name a few), which is a series of questions that address many different areas of data management, including context, technical specifications (formats, file types, sizes, software, etc.), policies, opinions, and needs. We meet with researchers to have a conversation, educate them on emerging trends and regulations in data sharing, and listen to their concerns and challenges. In the end, we try to make recommendations on how they can improve their data management processes, and then we offer to connect them with people who can help with the specific details (if it isn't us). For the DMPs, we have a series of templates that are specifically configured for the respective program requirements. Again in this case, we do some education, then offer some feedback and advising on what qualifies as good data management decisions for a particular community. Behind all of these efforts, we know we don't know all the answers, but we do know most of the questions to ask and who we need to pull together to figure out the solutions. That's our basic operating principle.

So, sound a bit familiar? Based on conversations with Tom, and reading some of the posts in the AIMS blog myself, it sounds like we are up against some very similar challenges in regards to the front-end of the issue, around education, conducting inventories and assessments, and figuring out how to manage processes before it comes down to managing the information itself and providing access to it for others. Appraisal and selection is incredibly important to us, but is usually driven more by the type of data. As an example, reproducible data generated by a big machine might not be important to keep, but the instructions and context in which it is generated would be invaluable. On the other hand, data from natural observations (ie. like climate data) would be critical to save. These considerations are not always apparent to the researcher, as they often think within the context of their work, rather than others. I would expect that the back-end is even more similar, as we are all ultimately dealing with bits and bytes, formats, standards, and figuring out how to decide what to keep and how to do it.

Lastly, for now, I also would like to mention that I had the opportunity to attend the annual Duke-Dartmouth Advisory Council meeting at the Fall CNI Forum several months ago.

http://www.dartmouth.edu/~vox/0708/0218/infomgmt.html

As you'll read, this project aims to bring together stakeholders from all areas of digital information across the institution, to talk about and plan in a collaborative and strategic way. They aim to tackle the challenges of management, technology, policy, and hardest, culture. I was incredibly impressed by the vision of this undertaking, and hope that we can continue to refine our efforts at developing a collaborative digital information management strategy as well. In practical terms, we all need to try and be attentive to how our effort plugs-in with others around the institution. The issue of digital information management is undoubtedly a very big one, and requires coordination and collaboration across many experts in order to appropriately treat the various bits that we encounter. Doing so will hopefully also provide us with the ability to bring best practices from one challenge to another.

--

Andrew is currently the Head of Strategic Data Initiatives and the Scientific Data Consulting Group at the UVA Library.

Contact info: Andrew Sallans, Email: als9q@virginia.edu, Twitter: asallans


Friday, 1 April 2011

Digital Collaboration Colloquium

On Tuesday I attended the Digital Collaboration Colloquium event in Sheffield organised to mark the end of the White Rose Libraries LIFE-SHARE Project.

The day included a number of talks about how institutions can collaborate including an interesting account of the Wales Higher Education Libraries Forum (WHELF) and experiences from the Victoria & Albert Museum. Although the majority of examples focussed on digitisation the principles and lessons learnt were all equally appropriate to a born-digital context.

As part of the day I presented a Pecha Kucha session on the AIMS project and some of the digital collaboration tools that we have found to be effective including Skype and GoogleDocs. In you are not familiar with this format it involves a presentation of 20 slides, changing automatically every 20 seconds and despite cutting the content quite heavily I still found myself chasing to keep-up with the changes. Other sessions looked at digitisation in-situ in a public setting – bringing behind the scenes in-front of the curtain, and other sessions on the Knitting patterns project at Southampton, the Addressing History project based at EDINA and the Yorkshire Playbills project.

The afternoon included a presentation form our hosts on the LIFE-SHARE project and their experiences of the collaboration continuum and a roundtable session that led to a good discussion between panel and audience. With alot covered in a relaxed and friendly atmosphere there was plenty of networking and I’m sure everybody took something from the day.

The presentations are available via slideshare

Friday, 18 March 2011

Personal Digital Archiving Conference 2011

I had the good fortune to attend the 2011 Personal Digital Archiving conference at the Internet Archive, along with other colleagues on the AIMS project, including Michael Forstrom from Yale and Michael Olson, Peter Chan, and Glynn Edwards from Stanford. The conference was exceptional, and had a great range of presentations ranging from those on fairly pragmatic topics to the highly theoretical. There are a number of other blogs with comprehensive notes on the conference, and the conference's organizers have already provided a detailed listing of those. Instead, I'd just like to focus on what I considered the highlights of the conference.
  • Cathy Marshall's keynote was excellent. I have seen her speak before, and she presented an survey of her ongoing research into personal digital archives.
  • Jeremy Leighton John presented on work undertaken since the Digital Lives project at the British Library.
  • Judith Zissman presented on "agile archiving", similar to agile development, wherein individuals can continually refine their archival practices.
  • Birkin Diana presented on how Brown University is working to make their institutional repository a space for personal materials, and strategies that allow users to work on adding metadata iteratively.
  • Daniel Reetz presented on his DIY Book Scanner project, but also brought in detailed technical analysis about how image sensors in digital cameras work and how our brains process image data.
  • Jason Zalinger introduced the notion of Gmail as a "story-world" and presented some prototype tools and games to help navigate that world.
  • Cal Lee presented on introducing education about digital forensics to the archival curriculum.
  • Kam Woods also presented on applying digital forensics to the archival profession.
  • Sam Meister presented on the complex ethics of using forensics in acquiring and processing the records from start-up companies.
In addition, I presented with Amelia Abreu on "archival sensemaking", which introduces the notion of personal digital archiving practice as an iterative, context-bound process.

Saturday, 12 March 2011

Processing Born Digital Materials Using AccessData FTK

To follow up with my previous blog entry on "Surprise Use of Forensic Software in Archives", I have prepared a YouTube video "Processing Born Digital Materials Using AccessData FTK". I hope this video can give people more details on how FTK is being used at Stanford University Libraries. Take a look and let me know what you think.

I would like to say a few words on discovery and access even though it is not the topic of the video. After we process the files in FTK, one way to delivery the files is to store them in a Fedora repository and let people access our Fedora repository using a web browser through Internet. We have developed an alpha version of this model using files from the Stephen Jay Gould collection. Another way to provide access to the files is to let people use FTK to access the files in our reading room. I will write about that later.

Hope you enjoy the video.

http://www.youtube.com/watch?v=hDAhbR8dyp8

Friday, 4 March 2011

File type categories with PRONOM and DROID

In order to assess a born digital accession, the AIMS digital archivists expressed a need for a report on the count of files grouped by type. The compact listing gives the archivist an overview that is difficult to visualize from a long listing. The category report supplements the full list of all files, and helps with a quick assessment after creation of a SIP via Rubymatica. (In a later post I’ll point out some reasons why pre-SIP assessment is often not practical with born digital.)

At the moment we have six categories. Below is a small example ingest:

Category summary for accession ingested files
data3
moving image1
other2
sound2
still image26
textual12
Total46


Some time ago we decided to exclusively use DROID as our file identification software. It works well to identify a broad variety of files, and is constantly being improved. We initially were using file identities from FITS, but the particular identity was highly variable. FITS gives a “best” identity based meta data returned by several utility programs. We wanted a consistent identification as opposed to some files being identified by DROID, some by the “file utility” and some by Jhove. We are currently using the DROID identification by pulling the DROID information out of the FITS xml for each file. This is easy and required very little change to Rubymatica.

PRONOM has the ability to have “classifications” via the XML element FormatTypes. However, there are a couple of issues. The first problem is that the PRONOM team is focused primarily on building new signatures (file identification configurations) and doesn’t have time to focus on low priority tasks such as categories. Second, the categories will almost certainly be somewhat different at each institution.

Happily I was able to create an easy-to-use web page to manage DROID categories. It only took one day to create this handy tool, and the tool is built-in to Rubymatica. The Rubymatica file listing report now has three sections: 1) overview using the categories 2) list of donor files in the ingest with the PRONOM PUID and human readable format name 3) the full list of all files (technical and donor) in the SIP.

This simple report seems anticlimactic, but processing born digital materials consists of many small details, which collectively can be a huge burden if not properly managed and automated. Adding this category feature to Rubymatica was a pleasant process, largely because the PRONOM data is open source, readily available, and delivered in a standard format (XML). My thanks and gratitude to the PRONOM people for their continuing work.

http://www.nationalarchives.gov.uk/PRONOM/Default.aspx

http://droid.sourceforge.net/

As I write this I notice that DROID v6 has just been released! The new version certainly includes a greatly expanded set of signatures (technical data for file identifications). We look forward to exploring all the new features.

Tuesday, 22 February 2011

Arrangement and Description of born-digital archives

For the last two months the Digital Archivists have been trying to define the requirements of a tool to enable archivists to arrange and describe born-digital archives. To do this we have stood-back and reviewed the traditional skills and processes and whether changes are required or appropriate to accommodate the particular issues surrounding born-digital archives.

The components we identified were as follows:
• Graphical User Interface – needs to be clean and easy to use
• Intellectual Arrangement - must be easy and instinctive for archivists to use
• Appraisal – born-digital archives need to be appraised as much as their paper predecessors
• Rights and Permissions – to enable the management of access to the born-digital archives and also to demonstrate to 3rd part depositors that the material is safe in your care
• Descriptive Metadata – a term we have been using to relate to description information and to explicitly distinguish this from the technical metadata about each file
• Import/Export functionality – to import/export data with other tools
• Reporting – to provide a range of "views" for managing the digital assets

Through a series of user stories and scenarios we have sought to clearly explain the requirement and how this might relate to other functionality.

This work has been under-taken predominantly through the use of GoogleDocs and created a document that we can all access and edit, create diagrams and include screenshots as necessary. Over weeks hundreds of comments have been added, and the text subjected to a comprehensive review and refinement process by numerous staff across the four partners.

Each institution has now scored and prioritised these features and as befits a collaborative initiative like the AIMS project allow us to identify a core group of features and functionality that we feel will be of greatest use to our institution and the wider archival community.

With the exception of intellectual arrangement most tasks and processes are not unique to archives so there is already a body of knowledge and experience in how to approach the task. For intellectual arrangement we have to be clear and precise about what we need and want we didn’t, for example a single intellectual arrangement when multiple versions would be possible in a digital environment.

Over the next few months we will be refining and reviewing these requirements, very much aware that there are only seven months of the project remaining. We also intend to discuss those aspects we identified as "critical" in future blog postings.

Tell us what tools you use with born-digital archives...

Friday, 19 November 2010

Surprise Use of Forensic Software in Archives

When I first heard of the use of computer forensic in archives, I was excited and wanted to learn how these law enforcement techniques could help me do a better job in processing digital collections. After learning people are using computer forensic to copy disk image (i.e. an exact copy of a disk, bit by bit) and to create a comprehensive manifest of the electronic files of collections, I was a bit disappointed because software engineers have been using the Unix dd command for many years to copy disk images. Also, there are tools (e.g. Karens's Directory Printer) available to create comprehensive manifest of the electronic files of collections. Data recovery is another feature of forensic software some people consider useful for archivists/researchers. In my opinion, data recovery may be useful for researchers but not archivists. Without informed written consents from donors, archivists should NOT recover the deleted files at all. Also, in some cases, a deleted file doesn't appear as one file, but instead, tens or hundreds of segments of files. When most archivists don't do item level processing in paper collections due to limited resources, I can't image archivists performing sub-item level processing in digital collections. Computer forensic in criminal application usually look for particular evidences. Organizing all files in a disk drive is usually not their interest. In archives, we are organizing all files in disk drives, looking for particular items are not our duty. Computer forensic may be more useful for researchers when they want to look for particular items. All these lead me to the conclusion that computer forensic may not be very useful for digital archivists.

However, after attending a 2.5-days training on AccessData FTK (a computer forensic software), I started to see the potential of using forensic software to process digital archives. I found out that the functions (bookmarks, labels) which help investigators to organize the evidence they selected are equally applicable to the organization of the whole collection. The functions (pattern and full text search) which are used to found particular evidence are equally applicable to search for restricted materials. I can also use the software to replace a number of software I am using to processing digital collections. Although, 90% of the training is related to cracking passwords, searching on delete files, identifying pornographic images, etc., I found the 10% the course worth every cents Stanford spent on it. Of course, the ideal case would be a course tailored for the archival community, but unfortunately, there is no such course exists.

Now, I am using AccessData FTK to replace the following software I used in the past to process digital archives.
Karens's Directory Printer - to create a comprehensive manifest of the electronic files of collections
QuickView Plus - to view files with obsolete file formats
Xplorer - to find duplicate files, copy to folders
DROID, JHOVE - to extract technical metadata: file formats, checksums, creation/modification/access dates
Windows Search 4.0 - to perform full text search on files with certain formats (word, pdf, ascii);

I am using following functions, which I have not found software package to perform in a very user friendly manner, in AccessData FTK to process digital archives.
Pattern search (to locate files containing restricted information such as social security no, credit card no., etc.)
Assign bookmarks, labels to files (for arranging files into series/subseries, other administrative and descriptive metadata)
Extract email headers (to; from; subject; date; cc/bcc) from emails written in different email programs for preparing correspondence listing.

The cost of licensing the software seems high. But if you look at the total costs of learning several "free" software, the lack of support for such software, and the integrated environment you get in using on software, you may find the total costs of using commercial forensic software is cheaper than using "free" software.