Tuesday, 21 June 2011
Photographing the digital: creating images of Hull University Archives’ digital media
Over the last few months I have been working with the AIMS team at Hull University. My role entails getting stuck into some practical processing of the born-digital collections in the Hull University Archives as well as planning aspects of digital preservation. A lot of our work so far has been to discover and document the material that we already hold in what we thought were purely paper collections and I have written a workflow for the discovery of these items and their preparation for ingest into Fedora. As part of this workflow we decided to photograph all of the removable media we currently have and create a process for photography of new deposits when they arrive.
Why bother?
By retaining photographs of the original media alongside content we will be able to provide an image of the appearance of the original media to researchers if they request it. For the foreseeable future we are storing the image files on a shared drive, but they will eventually be stored as an element of metadata with the digital files in our Fedora Repository. We will be dealing with large numbers of media items so need to ensure consistency in the way the media is photographed and information recorded from those images.
Process
Having not previously numbered the discs, we decided on a simple running number within each accession. Despite our familiarity with labelling paper material, it seemed more complicated with digital. Our conservator advised against sticking labels (even conservation grade) onto the plastic casing of a floppy or Amstrad disc. Though a specialist CD marker can be used to label CDs, we were reluctant to permanently mark the items! After a worryingly long thought process we decided to stick to the old faithful method of writing in pencil on the existing label or case.
I then started planning the process. Despite trying to anticipate the different elements of information to include for each media type, it was only trial runs photographing actual media that gave the full picture - i.e. that Amstrad discs have three aspects to photograph (Side A, Side B and the edge). Lots of seemingly trivial questions arose - like whether to photograph the case or whether to photograph a label if blank. Getting the process right from the start will save time in the long run.
We decided to create a ‘clapperboard’ to photograph with the items for a failsafe way to ensure easy identification. I decided on a reusable form printed on a transparency which we can label with a drywipe marker. Putting theory into practice needed several trial runs; after each one I adapted the form and the procedure.
In addition I wrote up detailed notes describing the procedure for each type of media we anticipate encountering. We worked out a sensible image quality – so to ensure legibility of the labels without clogging up our servers with unnecessarily large images. Once the photographs have been taken they are renamed and filed. We also maintain an inventory of the items and record the media and label information alongside it. This ensures that if we send items (like our Amstrad discs) away to a third party we can match them to our records when they return.
This process has been satisfying to complete and enables us to tick at least one thing off our to-do list. Anyone can get this part of the process completed – even for material which is stored on a shared drive, photography of the original media is a useful process.
Wednesday, 18 May 2011
AIMS: the UnConference

Not two full weeks into my new job as Digital Archivist at UVa on the AIMS grant, I rolled up my sleeves to facilitate and host an unconference with my fellow Digital Archivists. Our unconference would be two full days of discussions, demonstrations, lightning talks, and networking with digital archivists from around the globe. At first the thought was a little terrifying – I’m not even fully sure I know what this job is yet, how could I actually lead discussions on the salient topics? But my fears were baseless: all the unconference attendees were thoughtful, articulate, and lively participants. I learned much more from them than they probably did from me.
The unconference was held on the 13th and 14th of May at the Omni Hotel in Charlottesville. The 27 participants represented libraries, archives, museums, and digital humanities centers across the US, Canada, and the United Kingdom. Despite the differences in our institutions, backgrounds, and training, we learned that we not only shared similar challenges, but also the same hopes for collaboration and innovation.
The first day started off with a round of lightning talks. Each participant had 5 minutes to present a topic, project, problem or idea that they were interested in talking about. The variety in the talks was remarkable to me, traversing the breadth and depth of all that can be thought of as “born-digital” and the many processes involved in managing it. The lightning talks were also great way to get an introduction to each participant as well as their perspective or the particular issues they were dealing with in their institution. A brief outline of each of the talks is available on the AIMS Unconference Wiki.
Thursday, 5 May 2011
Workshop on "Using FTK Imager and AccessData FTK to Capture and Process Born Digital Materials"
The workshop covered the following:
FTK Imager – how to:
1. Download and install the software (free software - http://accessdata.com/support/adownloads).
2. Create a forensic image of an USB flash drive.
3. Create a logical image of the same flash drive.
AccessData FTK – how to:
1. Load an image – for this workshop we used a sampling from the Stephen Jay Gould papers.
2. View technical metadata generated by the software.
3. Arrange column settings to see specific file attribute (e.g. duplicate files).
4. Search for social security numbers using pattern search.
5. Test the full-text search function.
6. Flag files with sensitive information with "privileged" tag (such as those with social security numbers, etc.)
7. Use the bookmark feature for hierarchical information and apply it to groups of files (e.g. series, subseries, etc.)
8. Label groups of files with user defined labels (controlled vocabulary for computer storage media, document type suggested in the workshop, subject headings or access rights, etc.)
9. View files with specific bookmarks and labels.
Many incoming collections are hybrid collections – containing both analog and digital material. The digital component will become even greater as we move forward. Empowering all archivists to use a tool such as AccessData FTK to process the digital materials would be very useful.
Friday, 8 April 2011
Data Management Planning?
Following on Tom's generous invitation to write a post for the AIMS partner blog, I am finally getting around to doing so. Tom and I have been holding monthly discussions about our respective projects since sometime last summer, and have talked in great length about the commonalities between what my group (the Scientific Data Consulting Group) is dealing with in regards to research data management versus what the AIMS group is dealing with in terms of born-digital archive material.
We have found that there are many areas of similarity, and that we face many of the same challenges, although we approach the problem quite differently and of course have entirely different terminology given our relative perspectives.
To get started, I have a pretty good understanding of the born-digital problem set, but have not been keeping detailed notes on the workflows and solutions that the AIMS group has identified as best practices throughout the life of this project. My intention for this post is to share the issues that we are dealing with in research data management and try to make some suggestions around areas where there may be overlap and opportunities for great information sharing and collaboration.
Starting this past January 18, 2011, the National Science Foundation (NSF) put into effect a new implementation of their pre-existing data management planning requirement. This revision now requires that researchers submit a 2 page data management plan (DMP) that specifies the steps they will take to share the data that underlies their published results. This DMP will undergo formal peer-review, require reporting in interim/final reports, and all future proposals. In effect, what one says must then be done, or else one runs the risk of losing future funding opportunities or worse, losing all funding for the institution from that particular agency. Although this requirement is focused on data sharing, it isn't possible for such an initiative to succeed without first addressing a mass of other data management issues, ranging from technical, to policy, to cultural. As we often point out these days, it is far easier to improve the process of data management up-front, in the operational process phase, than it is to begin thinking about how to share the data at the end of the project. I would expect that those attacking the born-digital archive problem can fully relate.
Here in the Scientific Data Consulting (SciDaC) Group in the UVA Library, we have been collecting and developing our local set of data management best practices for some time and have served as advisors to researchers in both the research data management and DMP development areas (they are of course interrelated, but sometimes have different levels of urgency). In doing so, we have developed what we call a "data interview/assessment" (based in large part upon the visionary work of others, Purdue's Data Curation Profile and work from the UK's Digital Curation Centre, to name a few), which is a series of questions that address many different areas of data management, including context, technical specifications (formats, file types, sizes, software, etc.), policies, opinions, and needs. We meet with researchers to have a conversation, educate them on emerging trends and regulations in data sharing, and listen to their concerns and challenges. In the end, we try to make recommendations on how they can improve their data management processes, and then we offer to connect them with people who can help with the specific details (if it isn't us). For the DMPs, we have a series of templates that are specifically configured for the respective program requirements. Again in this case, we do some education, then offer some feedback and advising on what qualifies as good data management decisions for a particular community. Behind all of these efforts, we know we don't know all the answers, but we do know most of the questions to ask and who we need to pull together to figure out the solutions. That's our basic operating principle.
So, sound a bit familiar? Based on conversations with Tom, and reading some of the posts in the AIMS blog myself, it sounds like we are up against some very similar challenges in regards to the front-end of the issue, around education, conducting inventories and assessments, and figuring out how to manage processes before it comes down to managing the information itself and providing access to it for others. Appraisal and selection is incredibly important to us, but is usually driven more by the type of data. As an example, reproducible data generated by a big machine might not be important to keep, but the instructions and context in which it is generated would be invaluable. On the other hand, data from natural observations (ie. like climate data) would be critical to save. These considerations are not always apparent to the researcher, as they often think within the context of their work, rather than others. I would expect that the back-end is even more similar, as we are all ultimately dealing with bits and bytes, formats, standards, and figuring out how to decide what to keep and how to do it.
Lastly, for now, I also would like to mention that I had the opportunity to attend the annual Duke-Dartmouth Advisory Council meeting at the Fall CNI Forum several months ago.
http://www.dartmouth.edu/~vox/0708/0218/infomgmt.html
As you'll read, this project aims to bring together stakeholders from all areas of digital information across the institution, to talk about and plan in a collaborative and strategic way. They aim to tackle the challenges of management, technology, policy, and hardest, culture. I was incredibly impressed by the vision of this undertaking, and hope that we can continue to refine our efforts at developing a collaborative digital information management strategy as well. In practical terms, we all need to try and be attentive to how our effort plugs-in with others around the institution. The issue of digital information management is undoubtedly a very big one, and requires coordination and collaboration across many experts in order to appropriately treat the various bits that we encounter. Doing so will hopefully also provide us with the ability to bring best practices from one challenge to another.
--
Andrew is currently the Head of Strategic Data Initiatives and the Scientific Data Consulting Group at the UVA Library.
Contact info: Andrew Sallans, Email: als9q@virginia.edu, Twitter: asallansFriday, 1 April 2011
Digital Collaboration Colloquium
The day included a number of talks about how institutions can collaborate including an interesting account of the Wales Higher Education Libraries Forum (WHELF) and experiences from the Victoria & Albert Museum. Although the majority of examples focussed on digitisation the principles and lessons learnt were all equally appropriate to a born-digital context.
As part of the day I presented a Pecha Kucha session on the AIMS project and some of the digital collaboration tools that we have found to be effective including Skype and GoogleDocs. In you are not familiar with this format it involves a presentation of 20 slides, changing automatically every 20 seconds and despite cutting the content quite heavily I still found myself chasing to keep-up with the changes. Other sessions looked at digitisation in-situ in a public setting – bringing behind the scenes in-front of the curtain, and other sessions on the Knitting patterns project at Southampton, the Addressing History project based at EDINA and the Yorkshire Playbills project.
The afternoon included a presentation form our hosts on the LIFE-SHARE project and their experiences of the collaboration continuum and a roundtable session that led to a good discussion between panel and audience. With alot covered in a relaxed and friendly atmosphere there was plenty of networking and I’m sure everybody took something from the day.
The presentations are available via slideshare
Friday, 18 March 2011
Personal Digital Archiving Conference 2011
- Cathy Marshall's keynote was excellent. I have seen her speak before, and she presented an survey of her ongoing research into personal digital archives.
- Jeremy Leighton John presented on work undertaken since the Digital Lives project at the British Library.
- Judith Zissman presented on "agile archiving", similar to agile development, wherein individuals can continually refine their archival practices.
- Birkin Diana presented on how Brown University is working to make their institutional repository a space for personal materials, and strategies that allow users to work on adding metadata iteratively.
- Daniel Reetz presented on his DIY Book Scanner project, but also brought in detailed technical analysis about how image sensors in digital cameras work and how our brains process image data.
- Jason Zalinger introduced the notion of Gmail as a "story-world" and presented some prototype tools and games to help navigate that world.
- Cal Lee presented on introducing education about digital forensics to the archival curriculum.
- Kam Woods also presented on applying digital forensics to the archival profession.
- Sam Meister presented on the complex ethics of using forensics in acquiring and processing the records from start-up companies.
Saturday, 12 March 2011
Processing Born Digital Materials Using AccessData FTK
I would like to say a few words on discovery and access even though it is not the topic of the video. After we process the files in FTK, one way to delivery the files is to store them in a Fedora repository and let people access our Fedora repository using a web browser through Internet. We have developed an alpha version of this model using files from the Stephen Jay Gould collection. Another way to provide access to the files is to let people use FTK to access the files in our reading room. I will write about that later.
Hope you enjoy the video.
http://www.youtube.com/watch?v=hDAhbR8dyp8
