Tuesday, 17 April 2012

Workshop 1: for historians

Bookings for this event are now closed.


Could you imagine what historical questions you might be able to answer using a comprehensive archive of UK websites for the period 1996 to 2010 ? If so, this workshop may be for you, and bookings are now open.

The workshop affords a unique opportunity to learn about, and shape the development of, a unique new dataset, purchased by the JISC from the Internet Archive, and in the keeping of the British Library.

Where:  British Library (St Pancras, London)
When:  Thursday 24 May,  11am - 3.15pm

Sessions include:
  • Introducing the UK Web Archive 
  • How one historian might use the Domain Dark Archive  [see earlier post for a preview]
  • What is analytical access anyway, and what could it do for me ? 
To book a place, contact the project manager, Dr Peter Webster, at Peter.Webster@sas.ac.uk [or, from May 7, jane.winters@sas.ac.uk], with a brief statement of your research interests.
Booking is free, but places are very limited. Bookings will close at 12 noon on  May 10th, and applicants will hear whether they have secured a place soon afterwards. The event will be most suitable for scholars at doctoral level or higher.
Lunch will be served, and we will also be able to reimburse reasonable travel expenses within the UK.

Saturday, 14 April 2012

What on earth would I do with this data ?

We’re busy arranging a series of workshops in May and June. Their purpose is to gather humanities and social science scholars together to think collectively about the kind of purposes to which they might put a near-comprehensive dataset of the UK web domain 1996-2010.

The exercise is going to involve the use of the imagination to an extent, since part of this project is to help the British Library to design a user interface for the new dataset; so there isn’t yet anything ‘to play with’, as it were. In order to help fund scholars’ imaginations, I’ve started to sketch how I myself, as an historian of contemporary British Christianity, might start to use the dataset; what questions I would like to ask of it.

I come to this with a research interest in the forms of words in which religion (broadly defined) is discussed, and how those modes of discourse change over time. This can usefully be thought of using the following scheme:
(i) there are some perennial issues, in relation to what we might call constitutional Christianity, taking such questions as the position of the bishops in the House of Lords, and the establishment of the Church of England
(ii) there are older issues that have been ‘reactivated’ in recent years. For instance, denominational church schools were an issue as far back as the 1906 general election. After a period of calm about the issue in public discussion, the last decade or so has seen the issue come back to prominence - except, of course, that they are now known as faith schools.
(iii) there are also new issues, the obvious one being the perception of a threat from radical Islamism; an issue that was simply absent until relatively recently.

I personally am particularly interested in the domain dark archive, since the period 1996-2010 frames many of these issues perfectly. So, what might I ask of the archive, and which tools might I use ?

Basic visualisation: the Ngram

At a most basic level, I might want to look at the incidence of particular terms, and look for periods in which a particular term is employed more often. For this, there is the Ngram; a visualisation tool that is already employed by Google, and on the existing UK Web Archive. Consider the following case:  in February 2008, the archbishop of Canterbury Rowan Williams gave a lecture to an audience of lawyers which reflected on the scope for the incorporation of sharia law into UK law. For some details of the media storm that followed, see here. An Ngram of the incidence of the word 'sharia' in the existing selective web archive looks like this:
As we might expect, there is a big spike in the incidence of the term at the time of the lecture, and then heightened activity for much of the following year. I had expected the former, certainly, but not the latter to the same extent; and so I now know to look more at the content indicated by those subsequent spikes in activity.
If one then looks for both of the terms 'sharia' and 'archbishop', it appears:

The spike in the terms happens at roughly the same point; but the incidence of 'archbishop' is higher, due perhaps to the wider speculation about Dr Williams' position as a result of the controversy. Also, the repeated peaks visible for 'sharia' aren't present for 'archbishop', suggesting that the debate about the former outlasted the particular instance of the lecture.

Proximity searching and sentiment analysis

One might, of course, want to go further than this, and by means that aren't yet possible within the UK Web Archive. One means might be using a proximity search - looking for terms occurring within a certain number of characters' distance of each other in the same source.  The graph above only shows the instances of the two terms, but (crucially) not necessarily occurring together in the same source. A proximity search would make the connection that is suggested by the graphs above much more secure.

Even more interesting would be sentiment analysis: gauging the attitude of the writer of a webpage towards the term employed, using various techniques including natural language processing to find terms denoting approval or disapproval occurring in connection with the search term. The present archbishop, when he retires at the end of the year, may look back on a very particular relationship with the media during his time in office. I would be interested to see whether 'archbishop' appeared more often in the data with negative connotations after the sharia controversy.

These, of course, are only some attempts to imagine what might be possible using the Domain Dark Archive. I shall be blogging more as the project progresses, and the possibilities become clearer.



Wednesday, 14 March 2012

AADDA's 'elevator pitch'

I'm just on my way home from a very productive meeting with all the projects in this JISC sustainability strand. One of the activities during the day was building an one minute "elevator pitch" for the project, using the Pitch Builder from Harvard Business School. And so - here it is:

"AADDA is a joint venture between the IHR, the British Library and the University of Cambridge. It aims to transform the way in which humanities and social science researchers interact with the single most important archive of .uk web materials. It will develop innovative tools for analytical access to 40TB of primary data from UK webspace (1996-2010) and as a result will allow scholars to ask hitherto impossible questions of a singularly significant dataset. The archival record of contemporary Britain has increasingly migrated to a digital-only environment. The sheer volume of the record now requires new tools to render it accessible to scholars, and to unlock this unique and largely unexplored resource. In the next year, we will draw on a committed group of researchers to guide the British Library in the specification and development of tools for the analysis of the domain archive. Use cases arising from the project will be integral to the Library's sustainability strategy for the archive."

Thursday, 2 February 2012

The AADDA project


AADDA is an 18-month project to enhance the sustainability of a substantial dark archive of UK domain websites collected between 1996 and 2010 by the Internet Archive, copies of which were recently acquired by the JISC and are stored at the British Library on their behalf. The project team will work with researchers in contemporary history in particular, and digital humanities in general, to obtain feedback on the feasibility of using web archives at an analytical level. The project will build on this feedback and on the existing UK Web Archive interface in order to develop new forms of analytical access to this collection, thereby enabling researchers to carry out unique and hitherto impossible research queries. This will make a significant contribution to the global understanding of the research value of web archives, particularly for collections that span over a decade and more. Proven clarification of the utility of web archives for scholarly research will significantly enhance the long-term sustainability of the collection and provide valuable data about use cases for justification of ongoing funding.
The project will assess and aim to increase the acknowledged value of domain web archives for scholarly research. Following a survey of current perceptions and consultation with researchers, it will develop prototype tools for the exploitation of domain web archives, raise awareness of the material and services available, promote discussion and debate among key stakeholders, and inform future scholarly access arrangements at a domain level.
The project is led by Dr Jane Winters (IHR), and managed by Dr Peter Webster (IHR), working closely with Maureen Pennock and colleagues at the British Library and Dr Anne Alexander (University of Cambridge). Evaluation will be undertaken by Simon Tanner (King's College London).