Posts tonen met het label keywords. Alle posts tonen
Posts tonen met het label keywords. Alle posts tonen

dinsdag 20 oktober 2015

Inzoomen op de outliers.

Geregeld sprak ik hier over de trefwoorden en hun rol in trefwoordnetwerken. Ik liet grafiekjes zien van netwerken op basis van de brugfunctie die trefwoorden kunnen vervullen in netwerken (betweenness), over de veronderstelde invloeden van een trefwoord in een netwerk (eigenvector), over trefwoordmanifestaties in de OPC, Plinklets etc.

De gedachte is dat trefwoorden met een hoge betweenness en / of eigenvector waarde een -zeg- meer belangrijke rol spelen in het trefwoordnetwerk. Dit lijkt bevestigd te worden door het grove, oorzakelijke verband dat tussen beide waarden kan worden aangetoond. Zonder naar een beeld van een dergelijk netwerk te kijken weten wij al dat de geografische trefwoorden uit de aard der zaak een dergelijke rol zullen spelen. Dat komt, omdat dit soort trefwoorden eigenlijk overal kunnen opduiken: Nederland en piraterij, Nederland en familie recht, Nederland en terrorisme. Binnen een netwerk van aan elkaar gerelateerde onderwerpen vervuilen de geografische aanduidingen eigenlijk of, anders geformuleerd, zijn zij van een andere orde. In het navolgende heb ik daarom de geografische trefwoorden uitgefilterd. Bovendien beperk ik mijzelf in eerste instantie tot gegevens uit de maand augustus.

De vraag: "Welke zijn nu de trefwoorden die een relatief hoge betweenness en eigenvector waarde hebben?" is met behulp van de programma's Gephi en  R vrij eenvoudig te beantwoorden. Eerder zei ik al dat er een oorzakelijk verband is tussen de betweenness en eigenvector waarden: een hoge eigenvector waarde heeft bij hetzelfde trefwoord ook een hogere betweenness waarde en omgedraaid. Per trefwoord kunnen de verhoudingen overigens wel verschillen. Als je beide waarden in een grafiek uitzet dan zie je dus een denkbeeldige lijn tussen de trefwoorden door van grofweg linksonder naar rechtsboven. In de gegevens hieronder worden alleen hogere betweenness en eigenvector waarden meegenomen, maar niet de allerhoogste, die van Human rights of European Union bijvoorbeeld. Alle waarden meenemen levert een volledig volgelopen grafiek op, want dan duiken ook de trefwoorden op met wel heel lage waarden.


Links op de y-as zien we een niet realistische aanduiding van de getalswaarden. Ik heb de waarden opgerekt met een factor 8 om een betere vlakverdeling zichtbaar te maken. Op de x-as staat een wetenschappelijke notatie van hele lage eigenvector waarden. Deze waarden worden altijd in heel lage waarden aangeduid, vandaar. Iedere punt is een trefwoord.

Als ik nu in R instel dat we drie clusters moeten aanwijzen op basis van de eigenvector waarden dan resulteert dat in het volgende plaatje.



Programma R groepeert dus zoals getoond. Met het oog op eerste grafiek zou de mens wellicht meer clusters herkennen, maar ik heb heel expliciet aangegeven dat we met drie clusters werken. Toch is nu al te concluderen dat er meer trefwoorden met een lagere eigenvector- en betweennesswaarden zijn (zwart) dan trefwoorden met hogere waarden. Dat is niet verrassend natuurlijk. Als we hetzelfde doen, maar dan met betweennesswaarden dan ziet dat er zo uit.

Het vraagt een aparte studie om de clusters met elkaar te vergelijken en bijvoorbeeld eens te kijken naar de overlap in het groene cluster hierboven en het rode cluster in de grafiek daar weer boven. Maar het is natuurlijk ook mogelijk om te bekijken in hoeverre de clusters kleur krijgen als de aantallen manifestaties van de trefwoorden als leidraad voor de clustering worden genomen. En dat gebeurt hieronder.

De positie van het trefwoord in de grafiek wordt dus bepaald door de beide waarden, de kleur door het aantal. En dan is links onder ineens interessant, want het is vooral daar dat een zekere vermenging optreedt. Rood=300-1000, groen=1000-1250 en zwart=1250-3000 manifestaties. De groene stippen helemaal links, die gezien de aantallen manifestaties in relatie tot hun positie in de grafiek opvallend zijn, zijn de volgende trefwoorden (met * de trefwoorden die altijd opduiken) en hun aantallen:

Augustus:
Foreign direct investment - 1158, *Public international law - 1138, *Private International Law - 1127, *International criminal law - 1022, *International law - 1005
Wie had dat gedacht? Een relatief lage brugfunctie, een relatief lage invloed op de omgeving, waardoor kan dan de rol van het trefwoord 'Foreign direct investment' in de maand augustus 2015 worden verklaard?

De situatie in juli:
*International law - 1488, *Public international law - 1425, *United Nations - 1216

In juni:
*International criminal law - 960, *International humanitarian law -743, *Public international law - 739, Environmental protection - 626

In mei:
*International criminal law - 1047, *International humanitarian law - 906, *United Nations - 831, *International law - 795

In februari:
*International criminal law - 878, *International law - 794, *International humanitarian law  - 724, Terrorism - 608 (Hebdo?), *United Nations - 591, Environmental protection - 507

In januari duikt Terrorism net op in de grafiek en dat wordt doorgezet in februari (Hebdo?). Verder zien we in twee maanden het trefwoord 'Environmental protection' opduiken als opvallende manifestatie. Niet de belangrijkste, maar wel een belangrijke in het oog springende manifestatie.

De vraag is nu, is er op basis van bovenstaande een maandelijks 'belangstellingen profiel' samen te stellen of niet? En wat betekent dat dan voor de bibliotheek?

woensdag 3 december 2014

Presenting keywords. Eh? Which keywords? And how?

It is always hard to get started with bibliographic research. Especially for those patrons and scholars who realize it is important not to just search by some words using a very general index, but to use keywords (sometimes also called subject headings). These users know that a very dedicated group of librarians has thoroughly examined the publications added to their OPAC and has enriched them with keywords. And therein lies a problem, because one might ask: "how do these keywords look like?".

In my last couple of blogs I was focusing on how to get an idea of subject areas (huge and small). I used Gephi to create maps which could indeed give an impression of subjects and subject areas. For technical reasons, however, these maps dealt with only a relatively small set of data. In the latest published map I only incorporated about 4,500 exposed titles to our users in the reading room of the library. This looks like it is much, but in fact it is not. Therefore smaller subjects may not appear during this specific time frame or stay unnoticed. To get a more thorough impression of the subjects and the keywords used in the library one should use as large a set of information as possible. Luckily I have collected such a set, but before I come to that, I like to sketch a situation.

So, suppose some patron is looking for publications about biological warfare and the Security Council (we are after all the Peace Palace Library). Chances are (about 70% of our users do so) he uses the all words index in our OPAC and types 'biological warfare security council'. No hits. He then tries 'biological warfare' using the same index. 92 hits, which he quickly scans to look for what interests him (about 6 pages of title info). He then tries 'security council'. 1,416, hits which he does not scan, because it is to much. Now suppose he all of a sudden realizes he should also use the keyword index. Biological warfare, 213 hits. Security council, 2,741 hits. Combined (we just assume that he knows how to do this), 1 hit (a freely available PDF file containing references to other freely available texts and links to websites).

It is my opinion that libraries should bring to the forefront sets of keywords, all related to one general subject. Library users then just have to check these overviews in order to comprehend which keywords they should use in their research. In order to build up this set it is important to collect just the keywords which were, preferably during a longer time span, exposed or used by our patrons. This way we can be sure all relevant keywords can be collected. 

In January of this year I started to collect all the records, which were, in one way or another, seen or used by our patrons and visitors. This includes the records 'seen' by search robots, like the ones from Google. The database (we use for this MongoDB as a database system) contains now almost 7,500,000 records, but keep in mind that less then 10% of this number actually can be attributed to human beings. Each record contains, publication id, ip-number, time stamp and the keywords belonging to this publication. Given the size of the database it is possible to collect the really used or exposed keywords related to just one general subject.

I managed to create such a set of keywords as an example. They all were used in some combination with my 'main' keyword Biological and chemical weapons. The set contains a little bit more than 410 different keywords, some used quite a lot, others just a few times. 

So what remains is, to determine what kind of presentation to use to get a quick and thorough impression about specific keywords. I decided for now to use Tableau and to draw three diagrams with different colors. Each diagram indicates the relative amount of use and the keyword description, each next diagram present the keywords used less and less. So if you are looking for a publication about chemical warfare and genetic manipulation, after a peek at the diagrams below you will know what keywords to use to get to this information. (Curious? Look at this chapter: Terrorism in the Genomic Age / John Ellis, 2004.) 

And let's not forget serendipity.





donderdag 27 november 2014

Keywords. Collecting data.

In the last couple of weeks I blogged about keywords as they were displayed to the users of the OPAC in the library of the Peace Palace. I showed a couple of maps built with Gephi, some exhaustive, others very detailed.

But, how did I collect and adapt the data to be used by Gephi? I already mentioned "exposed keywords to the user" in an earlier blog. So to start with; what is the meaning of "exposed keywords"? I mean with this "keywords such as they occur in the presentation of the titles which were actually seen, perhaps even read, by the user". I' am interested in these keywords.

The enumeration of keywords in just one title can indeed be considered as a very small network. All these keywords are somehow linked to one another. Therefore, the first step is to gather all the presented titles and the second step is to collect all these small networks of keywords and then, lastly, to create one huge file which can be used by Gephi.

In the table below I give some examples of the file structure. In the left column you see five keywords (for Gephi they are nodes), each with a count of one, called 'use'. Underneath that you see the unique combinations of the keywords (for Gephi they are edges), also with a count, called 'weight'. The number after capital P is a unique keyword identifier. Gephi likes doing arithmetic with simple codes instead of -sometimes- long strings with weird characters in it. In the middle column you see the same, except now the keywords are from another title. In the third column both sets of keywords are combined. Take notice of the keyword 'Women', it occurs in both titles, therefore in the third column the 'use' is raised to two. At the bottom of each column you see the corresponding Gephi map.

nodedef>name VARCHAR,label VARCHAR, use INT
P076244229,"Middle East",1
P076242366,"Women",1
P076239810,"Islam",1
P076243265,"Family law",1
P07624234X,"Islamic law",1
edgedef>node1 VARCHAR,node2 VARCHAR,weight INT
P076244229,P076242366,1
P076244229,P076239810,1
P076244229,P076243265,1
P076244229,P07624234X,1
P076242366,P076239810,1
P076242366,P076243265,1
P076242366,P07624234X,1
P076239810,P076243265,1
P076239810,P07624234X,1
P076243265,P07624234X,1




















































nodedef>name VARCHAR,label VARCHAR, use INT
P076239519,"Refugees",1
P076242986,"Asylum",1
P076242366,"Women",1
P356258076,"Girls",1
P076256758,"Immigration",1
P241824206,"Convention relating to the Status of Refugees (Geneva, 28 July 1951)",1
P252290518,"Sex crimes",1
P255990391,"Gender",1
P332051005,"E-docs",1
edgedef>node1 VARCHAR,node2 VARCHAR,weight INT
P076239519,P076242986,1
P076239519,P076242366,1
P076239519,P356258076,1
P076239519,P076256758,1
P076239519,P241824206,1
P076239519,P252290518,1
P076239519,P255990391,1
P076239519,P332051005,1
P076242986,P076242366,1
P076242986,P356258076,1
P076242986,P076256758,1
P076242986,P241824206,1
P076242986,P252290518,1
P076242986,P255990391,1
P076242986,P332051005,1
P076242366,P356258076,1
P076242366,P076256758,1
P076242366,P241824206,1
P076242366,P252290518,1
P076242366,P255990391,1
P076242366,P332051005,1
P356258076,P076256758,1
P356258076,P241824206,1
P356258076,P252290518,1
P356258076,P255990391,1
P356258076,P332051005,1
P076256758,P241824206,1
P076256758,P252290518,1
P076256758,P255990391,1
P076256758,P332051005,1
P241824206,P252290518,1
P241824206,P255990391,1
P241824206,P332051005,1
P252290518,P255990391,1
P252290518,P332051005,1
P255990391,P332051005,1

















nodedef>name VARCHAR,label VARCHAR, use INT
P076244229,"Middle East",1
P076242366,"Women",2
P076239810,"Islam",1
P076243265,"Family law",1
P07624234X,"Islamic law",1
P076239519,"Refugees",1
P076242986,"Asylum",1
P356258076,"Girls",1
P076256758,"Immigration",1
P241824206,"Convention relating to the Status of Refugee (Geneva, 28 July 1951)",1
P252290518,"Sex crimes",1
P255990391,"Gender",1
P332051005,"E-docs",1
edgedef>node1 VARCHAR,node2 VARCHAR,weight INT
P076244229,P076242366,1
P076244229,P076239810,1
P076244229,P076243265,1
P076244229,P07624234X,1
P076242366,P076239810,1
P076242366,P076243265,1
P076242366,P07624234X,1
P076239810,P076243265,1
P076239810,P07624234X,1
P076243265,P07624234X,1
P076239519,P076242986,1
P076239519,P076242366,1
P076239519,P356258076,1
P076239519,P076256758,1
P076239519,P241824206,1
P076239519,P252290518,1
P076239519,P255990391,1
P076239519,P332051005,1
P076242986,P076242366,1
P076242986,P356258076,1
P076242986,P076256758,1
P076242986,P241824206,1
P076242986,P252290518,1
P076242986,P255990391,1
P076242986,P332051005,1
P076242366,P356258076,1
P076242366,P076256758,1
P076242366,P241824206,1
P076242366,P252290518,1
P076242366,P255990391,1
P076242366,P332051005,1
P356258076,P076256758,1
P356258076,P241824206,1
P356258076,P252290518,1
P356258076,P255990391,1
P356258076,P332051005,1
P076256758,P241824206,1
P076256758,P252290518,1
P076256758,P255990391,1
P076256758,P332051005,1
P241824206,P252290518,1
P241824206,P255990391,1
P241824206,P332051005,1
P252290518,P255990391,1
P252290518,P332051005,1
P255990391,P332051005,1





Of course I do not create these lengthy (15.000 lines and more) files by hand. I wrote a couple of crude PHP scripts to generate a crude file. This file I clean up with R and Microsoft Excel and the resulting file is ready to be used by Gephi. The scripts use a MongoDB collection, which contains all the logging of OPAC use in our readingroom. It is possible to detect 'exposed titles' (so also the keywords therein) in this logging.

To conclude. This is all very technical stuff and we may not expect our users to do this kind of research themselves, based on rough data provided by the library. However, some library staff members should certainly be able to do this. And then communicate about the results, using interesting maps for instance. Communicate to management about library collection issues, communicate to users about trends, communicate about almost lost niches in the collection, communicate about actual, important subcollections which can be used in updating dossiers, research guides, alerting systems, etc.

My other blogs about 'Gephi in libraries':

donderdag 20 november 2014

Just below the surface.

Using Gephi("an interactive visualization and exploration platform for all kinds of networks") to create unprocessed maps of exposed keywords to the user in the library of the Peace Palace, will result in an image in which a few huge subjects will dominate. These subjects are indicators of the core business of the library: Human Rights, European Union, United States of America, International Law and International Criminal Law to name just a few.To the left you see a very reduced image of such a map, but a few main keywords are still discernible.
These extra large topics veil the keywords just below. Zooming in will eventually bring you to the overshadowed keywords, but at a very deep level, so you will lose an overview of the structure. To the left we have zoomed in on an area clearly dominated by 'Human rights'. Now if I remove 'Human rights' from this cluster, Gephi will recalculate a lot of values, because one predominant element has been removed. After all, all keywords consitute one network. So the map gets a new shape. Especially, if all of the above mentioned subjects are removed and that is exactly what I have done. All the veiled keywords will float to the surface.
Let us now choose another criterion, in stead of the number of times a keyword occurs, to create a map. Gephi gives us a few other options, one of them is betweenness centrality.

Et voilĂ , after using the option 'rank parameter' in Gephi and choosing for betweenness centrality a new overall map appears, now with new highlighted nodes or keywords. Before zooming in, I will try to explain what betweenness centrality is. In brief, betweenness centrality is an indicator value for a key position. The higher the value the more important the role of the keyword. This value is calculated by counting the shortest paths between two keywords in our network. The keyword which appears the most times as being in between two different keywords, has the highest betweenness centrality value; these keywords are brokers or intermediaries. I used these values to create the map at the left.

After zooming in a little on the section of the map where the overall keyword 'Human Rights' used to be, a new picture arises. We see keywords like Children, Women and Family law, all of course related to Human rights and quite a few of them with a high key position or broker value. In short a new picture of related subjects emerges, indicating what the library of the Peace Palace could provide to its users.

By the way, the relations of keywords with a high betweenness centrality are not restricted to just one general subject. Between this kind of keywords there could be dense relations to other general subjects as well, see the image to the left.

Using Gephi maps not only give students and scholars a tool in hand to explore the collections of libraries, it also is a clear reminder of the necessity of using keywords to conduct efficient bibliographic research.


Those of you who would like to have the data file used in Gephi to create all the maps shown, do contact me at a.janson at ppl dot nl.

My other blogs about 'Gephi in libraries':

donderdag 13 november 2014

Keywords! Maps! Let's dive in.


Last week I blogged about maps and keywords: Library and user: one interest? I presented a few maps, created with Gephi, with which I tried to compare the activities of the library staff with the interests of OPAC users. I talked about general subjects like 'international criminal law', 'space debris' and things like that.

These maps can also be used to get a detailed picture, although I admit the presented maps are a bit difficult to read after zooming in. However, librarians can use Gephi itself to do detailed research in order to find out what our patrons are looking for.
See for example this image, clipped from the Gephi overview graph frame, which shows keywords all about art, trade and illegal activities in just a tiny section of the map. I think librarians can use such insights to better facilitate their users, especially if they detect returning patterns in searches during a longer period.

If librarians can 'translate' these insights in more relevant acquisitions, improvements in their research guides (in this case the Peace Palace Library, Cultural Heritage) or write specific blogs or tweets, I'am sure interested visitors will return to the library.

Of course users can manipulate the map with OPAC searches and focus on just one group (to the left you see the International Criminal Law group), but even one large group can be quite intimidating. Nevertheless those users who take some time can obtain a thorough knowledge about keywords grouped around one or two core subjects of the Peace Palace Library. Just start selecting a group using 'Group Selector' then click the largest bubble and check all the other keywords in the 'Information Pane'.

Librarians may use some of the more specific possibilities Gephi offers to look at maps in a very specific way. They may use for instance "Betweenness Centrality numbers" to look at 'broker' keywords, thus getting an idea about intermediaries. This knowledge too, I repeat, can be used to better respond to the needs of library users. I will write about this another time.

donderdag 6 november 2014

Library and user: one interest?

Quote: "But also interesting is, to see whether the library staff takes the interests of the patrons into account while acquiring documents for their collection? That is a subject for another blog."

Here I am referring to an earlier blogpost in which I tried to show what our users are looking for in the OPAC of the Peace Palace Library. In order to make this happen I focused on the use of our link resolver and presentation of a general subject in this link resolver. I used Tableau to create some graphs. 


However, the same thing can be done on the basis of the title descriptions which appeared on the screens in the reading room of the library after a succesful search. So, I collected all these titles and used all the keywords added to these titles to create a map using Gephi. In yet another blog I reported about this, although over there I used the recent acquisitions of the month of September.


In order to gain insight to answer the question "whether the library staff takes the interests of the patrons into account while acquiring documents for their collection?" I created two maps for comparison. One about the acquisitions in October and the other about the use of the OPAC in the readingroom in the same month. 



Acquisitions OPAC

If I enumerate the main subject topics which can be identified on indicated webpages, we get the following lists:


Acq:
  • International criminal law
  • Human rights
  • European Union
  • International trade
  • History
  • United Nations
  • Private international law
  • Islam/Islamic law
  • Law of the sea
  • Immigration
OPAC:
  • International criminal law*
  • Human rights*
  • European Union*
  • United Nations*
  • Intermational humanitarian law
  • International commercial arbitration
  • History/Politics*
  • Environmental protection
  • Law of the sea*
  • Space law

So our user behavior indicates special interest in Space, Environment, Commerce -among other things- which were not covered by our library staff. However, the library acquired material about Immigration, Islamic Law and Trade which was not looked for by our OPAC users. But of great importance is still the observation that both parties share their interest in the core business of the library of the Peace Palace: Criminal law, Human rights, European Union.

Only with regard to the peripheral areas differences exist and for a large part that can be related to current events, like boat refugees in the Mediterranean Sea, terrorism in the Middle East, space debris and environmental issues. 

Anyway, the simple fact that the 'small subjects' are also found and acquired, means that the library of the Peace Palace is on the right track. The 'small subjects' looked for now, were added in the past!


woensdag 29 oktober 2014

Subjects by keywords.

Each library buys books, journals or access to the e-version of these, files, databases, etc. In each library, there is a specific focus on a particular field of interest or -more often- multiple fields of interest. In the library of the Peace Palace 'documents' are acquired in the field of international law. Of course it is possible to recognize a large number of sub fields within this vast subject: international criminal law, human rights, diplomacy and so on.

For the users of the library, it is important to know which of these areas are covered and whether therefore it is worthwhile to use that library when you yourself are dealing with such a subject. One of the tools the library is using, is a system of sending regular 'alerts'. Interested scholars, students and other interested parties can be informed about the most recent acquisitions, once a week. There are around 1.000 subscribers to this service. They receive a weekly overview, based on a single criterion. The consequence of using a single general criterion is of course that the outcome may be huge. Especially topics like 'European Union' or 'Public International Law' may contain quite a lot of bibliographical references.

But what if you have a way of presenting the acquisitions in one month, not on the basis of one single general criterion, but where combinations of keywords assigned to each title play a role? What if relationships between those keywords can be visualized on the website in stead of emailed to a subscriber? In an attempt to make that possible I created a clickable map, which can be found on http://www.ppl.nl/september.

To create this map I used the visualization tool 'Gephi', which is especially strong in showing the links between the building blocks of the map. So I consider, for this purpose, the keywords as building blocks. On above mentioned web address, the acquisitions from the month of September are recorded, not in the form of boring title lists, but in the form of assigned keywords and relationships between those keywords. It is still all about numbers, as the strength of a relation is determined by the amount of occurrences of keyword pairs.

Of course in a batch of thousands of titles some areas in the map should be indicated by large blobs of tightly connected keywords. These blobs refer to the core businesses of the library. At the edges of the map, smaller subareas appear. In other words, large areas of the map show acquisitions that are always extremely important to the library. They can be recognized, because the largest circles in the map appear over there, surrounded by a large amount of closely packed smaller circles.


Subtopics are located outside the center of the map and can be considered as subjects which are farther away from the core business. So the area in the upper right is characterized by the keyword 'History', especially the history of the First World War. Typical related keywords are: 'Military History', 'Massacres', 'Ethnic Minorities', 'Russian Empire', etc. In the bottom left there is an area that relates to commerce, trade, international commercial arbitration, etc.

It must be stressed that subareas differ each month. It all depends on what is going on in the world. Highlighted in the last couple of months of this year is of course the commemoration of the start of the First World War. But other current events can also lead to a temporary increase in attention, such as transboundary pollution of the environment, disease outbreak, cyber warfare, sporting events, etc.

Clicking on a circle produces an 'Information Pane' where additional information can be found on that keyword, like how many times it occurs, but also the related keywords -those that are used in combination with the clicked keyword- are mentioned. So, with a few clicks it is easy to get a good impression about different topics and topic areas.

Finally, you are invited to use the map by hovering over the various items and view the information in the right pane. Are you looking for a specific topic? You can search by using the Search box in the left panel or zoom in on a specific group or cluster by 'Group Selector'.