Category Archives: DH Tools

Ekphrasis as an LDA Network in NodeXL

In an earlier post, I mention the value of visualizations as a means for exploring topic modeling data.Â That particular example used a small model of 276 poems labeled â€œekphrasticâ€ out of a much larger collection.Â At that point, I was still struggling with how to read the data, which felt overwhelming.Â How could I organize the relationships between topics and documents in such a way as to see salient connections produced by the model? Â The intermediate solution was to break the model down into groups of 3 topics and create bar graphs charting the likelihood that each document contained language from each topic.Â That solution worked in the short-term, because it helped me to discover the fact that one topic was found highly likely within a particular volume of ekphrastic verse: John Hollanderâ€™s The Gazerâ€™s Spirit.

Still, what I wanted was an impressionistic overview of the documentsâ€™ association with all of the topics. The first 40 or so attempts at this process were a dismal failure.Â Partly because it was a learning process and partly because the results frequently resembled the much maligned â€œhairball,â€ what I produced was completely incomprehensible.Â However, August 20th to 24th I attended the NSF, Social Media Research Foundation, and Grand funded Summer Social Webshop on Technology-Mediated Social Participation.Â There, I met Marc Smith, who began developing NodeXL, a social media network analysis tool built to work with Microsoft Excel, while he worked for Microsoft Research.Â Marc, who now leads the Social Media Research Foundation and Connected Action Â generously took time to demonstrate how to import my topic modeling data into NodeXL so that I could generate graphs that are more elegant and streamlined than any Iâ€™ve been able to produce to this point.Â The results arenâ€™t just beautiful: theyâ€™re useful.

So, what are those results? They include unimodal and bimodal network graphs that visualize connections between documents with other documents, topics with other topics, and documents with topics created with an LDA model in MALLET.Â Using NodeXLâ€™s algorithms, I am able to cluster groups with stronger ties in grid areas, assign them unique colors, and demonstrate the degree of probability the model calculates as a connection between nodes (either documents or topics depending on the graph).Â The real power of NodeXL, though, is that in the future I can make my data public through the NodeXL gallery, and you can download my network graph and play with it yourself.Â The data isnâ€™t quite there yet, but thatâ€™s whatâ€™s coming.

In the meantime, Iâ€™ll offer the following image of a network graph that I had hoped to produce with my earlier post about The Gazerâ€™s Spirit.Â Though the topic label is small, Topic 3 can be seen in the top left hand corner of the network diagram. The width and color of the edges in the diagram (meaning the width of the lines) is determined by the modelâ€™s estimation of how much of each topic is in each poem.Â If the lines are thicker and lighter, it means that the model estimates that a large portion of the poem draws its language from the corresponding topic.Â Similarly, the thinner and darker a line is the lower the probability that the poem includes language from the corresponding topic.

Table 1: Ekphrastic Dataset – 276 poems and 15 topics

Â Â Â Â Â Â Â Â Â Â Â Topic 3 (in the top, left-hand corner) is primarily comprised of connections to poems from The Gazerâ€™s Spirit and is affiliated by language that reflects a kind of courtship, including archaic references (thy, thee, thou) and the language of love (er, beauty, grace, eyes, heaven, divine, hand, love).Â This makes sense in the context of existing knowledge about Hollanderâ€™s volume.Â The collection reads very much like a tribute to painting and the visual arts by poetry, and the language of desire is prevalent throughout.Â Moreover, both W.J.T. Mitchell and James A.W. Heffernan, two prominent theorists in the ekphrastic tradition, insist that the language of love and desire is a strong, if not dominant, discourse across all of ekphrasis based on a canon of poems mostly included in The Gazer’s Spirit.Â One might assume, then, that there would be strong connections between a topic comprised of the language of courtship, love, and desire and most of the poems in the collection; however, only a few of the poems with a statistically significant portion of its language from Topic 3 are not also in The Gazerâ€™s Spirit: â€œThe Picture of Little T.C. in a Prospect of Flowers,â€ â€œThe Art of Poetry [excerpt],â€ â€œOzymandius,â€ â€œCanto I,â€ and â€œMy Last Duchess.â€Â Of those poems, none are by female poets.

Poems with highest proportion of Topic 3

The Temeraire (Supposed to Have Been Suggested to an Englishman of the Old Order by the Flight of the Monitor and Merrimac) by Herman Melville

To my Worthy Friend Mr. Peter Lilly: on that Excellent Picture of His majesty, and the Duke of York, drawne by him at Hampton-Court by Sir Richard Lovelace

From The Testament of Beauty, Book III by Robert Bridges

For Spring By Sandro Botticelli (In the Academia of Florence) by Dante Gabriel Rosetti

To the Statue on the Capitol: Looking Eastward at Dawn by John James Piatt

The Poem of Jacobus Sadoletus on the Statue of Laocoon by Jacobus Sadoleto

To the Fragment of a Statue of Hercules, Commonly Called the Torso by Samuel Rogers

The Last of England by Ford Maddox Brown

On the Group of the Three Angels Before the Tent of Abraham, by Rafaelle, in the Vatican by Washington Allston

Death’s Valley To accompany a picture; by request.Â “The Valley of the Shadow of Death,” from the painting by George Inness by Walt Whitman

Elegiac Stanzas Suggested by a Picture of Peele Castle, in a Storm, Painted by Sir George Beaumont by William Wordsworth

On the Medusa of Leonardo da Vinci in the Florentine Gallery by Percy B. Shelley

The Mind of the Frontispiece to a Book by Ben Jonson

Venus de Milo by Charles-Rene Marie Leconte de Lisle

The City of Dreadful Night by James Thomson

Sonnet by Pietro Aretino

For “Our Lady of the Rocks” By Leonardo da Vinci by Dante Gabriel Rosetti

Mona Lisa by Edith Wharton

Ode on a Grecian Urn by John Keats

The National Painting by Joseph Rodman Drake

The “Moses” of Michael Angelo by Robert Browning

Hiram Powers’ Greek Slave by Elizabeth Barrett Browning

From Childe Harold’s Pilgrimage, canto 4 by George Byron Gordon

The Picture of Little T. C. in a Prospect of Flowers by Andrew Marvell

Before the Mirror (Verses written under a Picture)Inscribed to J. A. Whistler by Algernon Charles Swinburne

For Venetian Pastoral By Giorgone (In the Louvre) by Dante Gabriel Rosetti

The Art of Poetry [excerpt] by Nicolas Boileau-Despreaux

Ozymandias by Percy B. Shelley

The Iliad, Book XVIII, [The Shield of Achilles] by Homer

Canto I by Dante Alighieri

The Hunter in the Snow by William Carlos Williams

Tiepolo’s Hound by Derek Wallcot

St. Eustace by Derek Mahon

Three for the Mona Lisa by John Stone

My Last Duchess by Robert Browning

Table 2: Ekphrastic Dataset 15 Topic Model, Topic 3 Highlighted

Â The only remaining topic which includes the word love fairly high in the key word distribution is Topic 4, which includes the following terms: portrait, monument, foreman, felt, woman, monuments, box, press, bacall, detail, young, thick, crimson, instrument, hotel, compartment, picked, cornell, Europe, lovers. As you can see from the network diagram below, none of the topics with high probabilities of containing Topic 3 are included in the Topic 4 distribution.

Table 3: Ekphrastic Dataset 15 Topic Model, Topic 4 Highlighted

Equally interesting, poems with the highest proportion of Topic 4 are also authored by female poets. Â Certainly, more poems by men include significant proportions of Topic 4 than poems by women that include significant portions of Topic three; however, there are striking and salient points to be made about the contrasting networks:

Poems with highest proportion of Topic 4

“Utopia Parkway” after Joseph Cornell’s Penny Arcade Portrait of Lauren Bacall, 1945 â€“ 46 by Linda Hull

Canvas and Mirror by Evie Shockley

Portrait of Madame Monet on Her Deathbed by Mary Rose Oâ€™Reilley

Internal Monument by G. C. Waldrup

The Uses of Distortion by Caroline Crumpacker

Joseph Cornell, with Box by Michael DumanisÂ Â

Drawing Wildflowers by Jorie Graham

The Eye Like a Strange Balloon Mounts Toward Infinity by Mary Jo Bang

Visiting the Wise Men in Cologne by J.P. White

Rhyme by Robert Pinksy

The Street by Stephen Dobyns

The Portrait by Stanley Kunitz

“Picture of a 23-Year-Old Painted by His Friend of the Same Age, an Amateur” by C.P. Cavafy

Portrait in Georgia by Jean Toomer

For the Poem Paterson [1. Detail] William Carlos Williams

The Dance by William Carlos Williams

Late Self-Portrait by Rembrandt by Jane Hirshfield

Sea Life in St. Mark’s Square by Mary Oâ€™Donnell

Washington’s Monument, February, 1885 by Walt Whitman

Still Life by Jorie Graham

Still Life by Tony Hoagland

The Family Photograph by Vona Groarke

The Corn Harvest by William Carlos Williams

Portrait of a Lady by T. S. Eliot

Portrait d’une Femme by Ezra Pound

This impressionistic overview of the ekphrastic dataset prompted through the exploration of a network graph of the relationships between topics and poems is a first step.Â Enough, perhaps, to formulate a new hypothesis about the difference between â€œloveâ€ and â€œloversâ€ in ekphrastic poetry, or to lend further support to the growing sense that there is a much broader range of kinds of attraction and kinshipâ€”a range inclusive of both competitive and kindred discoursesâ€”than previous theorizations of the genre have taken into account. Â The network visualization goes further than to suggest that there are two very different discourses regarding love and affection in ekphrastic verse, but even suggests possible poems to consider reading closely to see what those differences might be and if they are worth pursuing further. Â Through the use of networked relationships between topics and documents, we begin with lists of poems in which the discourse of affinity, affection, and desireâ€”as courtship or as partnershipâ€”can be further explored through close readings.

Meeting Edward Tufte’s claim that evidence should be both beautiful and useful, the NodeXL network diagrams of LDA data are a step toward developing methods of evaluating and exploring models of figurative language that do not necessarily fit the same criteria for models of non-figurative texts.

Why use visualizations to study poetry?

14 Replies

[Note: This post was a DHNow Editor’s Choice on May 1, 2012.]

The research I am doing presently uses visualizations to show latent patterns that may be detected in a set of poems using computational tools, such as topic modeling.Â In particular, Iâ€™m looking at poetry that takes visual art as its subject, a genre called ekphrasis, in an attempt to distinguish the types of language poets tend to invoke when creating a verbal art that responds to a visual one.Â Studying wordsâ€™ relationships to images and then creating more images to represent those patterns calls to mind a longstanding contest between modes of representationâ€”which one represents information â€œbetterâ€?Â Since my research is dedicated to revealing the potential for collaborative and kindred relationships between modes of representationÂ historicallyÂ seen in competition with one another, using images to further demonstrate patterns of language might be seen as counter-productive.Â Why use images to make literary arguments? Do images tell us something â€œnewâ€ that words cannot?

Without answering that question, Iâ€™d like instead to present an instance of when using images (visualizations of data) to â€œseeâ€ language led to an improved understanding of the kinds of questions we might ask and the types of answers we might want to look for that wouldnâ€™t have been possible had we not seen them differentlyâ€”through graphical array.

Currently, Iâ€™m using a tool called MALLET to create a model of the possible â€œtopicsâ€ found in a set of 276 ekphrastic poems.Â There are already several excellent explanations of what topic modeling is and how it works (many thanks to Matt Jockers, Ted Underwood, and Scott WeingartÂ who posted these explanations with humanists in mind), so Iâ€™m not going to spend time explaining what the tool does here; however, I will say that working with a set of 276 poems is atypical.Â Topic modeling was designed to work on millions of words, and 276 poems doesnâ€™t even come close; however, part of the project has been to determine a threshold at which we can get meaningful results from a small dataset.Â So, this particular experiment is playing with the lower thresholds of the toolâ€™s usefulness.

When you run a topic model (train-topics) in MALLET, you tell the program how many topics to create, and when the model runs, it can output a variety of results. Â As part of the tinkering process, Iâ€™ve been working with the number of topics to have MALLET use in order to generate the model, and was just about to despair that the real tests I wanted to run wouldnâ€™t be possible at 276 poems. Â Perhaps it was just too few poems to find recognizable patterns. Â For each topic assignment, MALLET assigns an ID number to the topic and “topic keys” as keywords for that topic. Â Usually, when the topic model is working, the results are â€œreadableâ€ because they represent similar language. Â MALLET would not call a topic “Sea,” for example, but might instead provide the following keywords:

blue, water, waves, sea, surface, turn, green, ship, sail, sailor, drown

The researcher would look at those terms and think, â€œOh, clearly thatâ€™s a nautical/sea/sailingâ€ topic, and dub it as such.Â My results, however, on 15 topics over 276 poems were not readable in the same way.Â For example, topic 3 included the following topic keys:

3Â Â Â Â Â Â Â Â Â 0.04026Â Â Â Â Â Â Â Â Â Â with self portrait him god how made shape give thing centuries image more world dread he lands down back protest shaped dream upon will rulers lords slave gazes hoe future

I donâ€™t blame you if you donâ€™t see the pattern there.Â I didnâ€™t.Â Except, well, knowing some of the poems in the set pretty well, I know that it put together â€œLandscape with the Fall of Icarusâ€ by W.C. Williams with â€œThe Poem of Jacobus Sadoletus on the Statue of Laocoonâ€ with â€œThe New Colossusâ€ with â€œThe Man with the Hoe Written after Seeing the Painting by Millet.â€Â I could see that we had lots of kinds of gods represented, farming, and statues, but thatâ€™s only because I knew the poems. Â Without topic modeling, I might put this category together as a â€œmastersâ€ grouping, but itâ€™s not likely. Â Rather than look for connections, I was focused on the fact that the topic keys didnâ€™t make a strong case for their being placed together, and other categories seemed similarly opaque.Â However, just to be sure that I could, in fact, visualize results of future tests, I went ahead and imported the topic associations by file.Â In other words, MALLET can also produce a file that lists each topic (0-14 in this case) with each file name in the dataset and a percentage.Â The percentage represents the degree to which the topic is represented inside each file.Â I imported the MALLET output of topics and files associated with them into Google Fusion Tables and created a dynamic bar graph that collects file-ids along the vertical axis and along the horizontal axis can be found the degree that the given topic (in this case topic 3) is present in the file.Â Â As I clicked through each topicâ€™s graph, I figured I was seeing results that demonstrated MALLETâ€™s confusion, since the dataset was so small.Â But then I saw this: [Below should be a Google Visualization. Â You may need to “refresh” your browser page to see it. Â If you still cannot see it, a static version of the file is visible here.]

If the graphâ€™s visualization is working, when you pass your mouse over the lines in the bar graph, the ones that are higher than 0.4, then the file-id number (a random number assigned during the course of preparing the data) appears. Â Each of these files begin with the same prefix: GS. Â In my dataset, that means that the files with the highest representation of topic 3 in them can all be found in John Hollanderâ€™s collection The Gazerâ€™s Spirit.Â This anthology is considered to be one of the most authoritative and diverseâ€”beginning with classical ekphrasis all the way up to and including poems from the 1980s and 1990s.Â I had expected, given the disparity in time periods, that the poems from this collection would be the most difficult to group together because the diction of the poems changes dramatically from the beginning of the volume to the end.Â In other words, I would have expected the poems to blend with the other ekphrastic poems throughout the dataset more in terms of their similar diction than by anything else.Â MALLET has no way of knowing that these files are included in the same anthology.Â All of the bibliographical information about the poems has been stripped from the text being tested.Â There has to be something else.Â What something else might be requires another layer of interpretation.Â I will need to return to the topic model to see if a similar pattern is present when I use Â other numbers of topicsâ€”or if I include some non-ekphrastic poems to the set being testedâ€”but seeing the affinity in language between the poems included in The Gazerâ€™s Spirit in contrast to other ekphrastic poems proved useful. Â Now, Iâ€™m not inclined to throw the whole test away, but instead to perform more tests to see if this pattern emerges again in other circumstances.Â Iâ€™m not at square one. Iâ€™m at a square 2 that I didnâ€™t expect.

The visualization in the end didnâ€™t produce â€œnew knowledge.â€Â It isnâ€™t hard to imagine that an editor would choose poems that construct a particular argument about what â€œbestâ€ represents a particular genre of poetry; however, if these poems did truly represent the diversity of ekphrastic verse, wouldnâ€™t we see other poems also highly associated with a â€œGazerâ€™s Spirit topicâ€?Â What makes these poems stand out so clearly from others of their kind?Â Might their similarity mark a reason for why critics of the 90s and 2000s define the tropes, canons, and traditions of ekphrasis in a particular vein?Â Iâ€™m now returning to the test and to the texts to see what answers might exist there that I and others have missed as close readers.Â Could we, for instance, run an analysis that determines how closely other kinds of ekphrasis are associated with Gazerâ€™s Spiritâ€™s definition of ekphrasis?Â Is it possible that poetry by male poets is more frequently associated with that strain of ekphrastic discourse than poetry by female poets?

This particular visualization doesnâ€™t make an “argument” in the way humanists are accustomed to making them.Â It doesnâ€™t necessarily produce anything wholly â€œnewâ€ that couldnâ€™t have been discovered some other way; however, it did help this researcher get past a particular kind of blindness and helped me to see alternativesâ€”to consider what has been missed along the wayâ€”and there is, and will be, something new in that.

Chunks, Topics, and Themes in LDA

4 Replies

[NB: This post is the continuation of a conversation begun on Ted Underwoodâ€™s blog under the post â€œA touching detail produced by LDAâ€â€”in which he demonstrates that there is an overlay between the works of the Shelley/Godwin family and a topic which includes the terms mind / heart / felt.Â Rather than hijack his post, Iâ€™m responding here to questions having to do more with process than content; however, to understand fully the genesis of this conversation, I encourage you to read Tedâ€™s post and the comments there first. ]

Ted-

I appreciate your response because it is making me think carefully about what I understand LDA “topics” to represent. Â Iâ€™m not sure that Iâ€™m on board with thinking of topics in terms of discourse or necessarily â€œwaysâ€ of writing. Â Honestly, Iâ€™m not trying to be difficult here; rather, Iâ€™m trying to parse for myself what I mean when I talk about my expectations that particular terms â€œshouldâ€ form the basis for a highly probable topic.Â It seems to me that what one wants from topic modeling are lexical themesâ€”in other words, lexical trends over the course of particular chunks of text.Â Iâ€™m taking to heart here Matt Jockersâ€™s recent post on the LDA buffet in which he articulates the assumption that LDA analysis makesâ€”that the world is composed of a certain number of topics (and in Mallet, we define those topics when we run the topic modeling application).Â As a result, when I run a topic model analysis in Mallet, I am looking at the way graphemes (because the written symbol, of course, is divorced from its meaning) relate to other similar graphemes.Â So, though topics may not have a one-to-one semantic relationship with particular volumes as the â€œmain topicâ€ or â€œsupporting topics,â€ one might reasonably expect that a text with a 90% probability of including a list of graphemes from an LDA topic lexicon (for lack of a better word) would correspondingly address a thematic topic which depends heavily on a closely related vocabulary.Â Similarly, the frequent use of words in a topic lexicon increases the probability that the LDA topic, through the repetition of those words, carries semantic weightâ€”though the degree to which this is the case wouldnâ€™t likely be determined by that initial topic probability.

Iâ€™m chasing the rabbit down a hole here, but I do so for the purpose of agreeing with your earlier claim that what kinds of results we get, their reliability, and their usefulness seems to be largely determined by the kinds of questions weâ€™re asking in the first place.Â I agree that when we use LDA to describe texts, thatâ€™s fundamentally different from using it to test assumptions/expectations.Â In my research, I have attempted to draw very clear distinctions between when I am testing assumptions about the kinds of language that dominate a particular genre of poetry and when I am using LDA to generate a list of potential word groups that could then be used to describe poetic trends.Â I see those as two very different projects.Â When Iâ€™m working with poetry and specifically with ekphrasis, I am testing what people who write about this particular genre assume to be true: that the word or variations of the word still will be one of the most commonly used words across all ekphrastic texts and used at a higher rate than in any other genre of poetry. Itâ€™s true that the word still could be a semantic topic in many other kinds of poetry; however, what weâ€™re trying to get at is that a group of words closely allied with the word still will be the most dominant and recurring trend across all ekphrastic verse.Â The next determination, then, to be made is whether or not that discovery carries semantic weight.Â If still, stillness, death, breathless, etc are not actually a dominant trend, have we overstated the case?

It seems that what youâ€™re saying (and please intervene if Iâ€™m not articulating this correctly) , which I tend to agree with is that â€œchunk sizeâ€ should be something determined by the questions being asked, and stating the way in which data has been chunked reflects the types of results we want to get in return.Â Taking this into consideration, though, certainly has helped the way I position what Iâ€™m doing.Â For me it is significant to chunk at the level of individual poems; however, were I to change my question to something like, â€œWhich poets trend more toward ekphrastic topics than others?â€â€”based on what weâ€™re saying here, that question seems to require chunking volumes rather than individual poems.

In other news, test models on the whole 4500 poems in my dataset, which is chunked at the level of individual poem, yielded much more promising initial results than we thought we would get.Â I would guess that it has something to do with the number of topics we assign when we run the model, and maybe one of the other ways forward is to talk about the threshold number of topics we need to assign in order to garner meaningful results from the model. Â (Obviously people like Matt and Travis have hands-on experience with this; however, I’m wondering if the type of question we’re asking should have a definable impact on how many topics we generate for the different types of tests….) Hopefully, in the near future Iâ€™ll be able to share some of those very preliminary resultsâ€¦ but Iâ€™m still in the midst of refining my queries and configuring my data.

Again, Iâ€™m engaged because I find what youâ€™re doing both relevant and useful, and I think that having these mid-investigation conversations does help to inform the way ahead.Â As you mention, perhaps many of these kinds of questions are answered in Matt Jockersâ€™s book, but it is unlikely Iâ€™ll be able to use that before this first iteration of my project is done in the next month or two. Â I believe that hearing anecdotal conversation about the low-level kinds of tests people are playing with really does help others along in their own work since we’re still figuring out what exactly we can do with this tool.

Small Projects & Limited Datasets

5 Replies

Iâ€™ve been thinking a lot lately about the significance of small projects in an increasingly large-scale DH environment.Â We seem almost inherently to know the value of â€œbig data:â€ scale changes the name of the game.Â Still, what about the smaller universes of projects with minimal budgets, fewer collaborators, and limited scopes, which also have large ambitions about what can be done using the digital resources we have on hand? Â Rather than detracting from the import of big data projects, I, like Natalie Houston, am wondering what small projects offer the field and whether those potential outcomes are relevant and useful both in and of themselves as well as beneficial to large-scale projects, such as in fine-tuning initial results.

My project in its current iteration involves a limited dataset of about 4500 poems and challenges rudimentary assumptions about a particular genre of poetry called ekphrasisâ€”poems regarding the visual arts.Â It is the capstone project to a dissertation in which I use the methods of social network analysis to explore socially-inscribed relationships between visual and verbal media and in which the results of my analysis are rendered visually to demonstrate the versatility and flexibility available to female poets writing ekphrastic poetry. My MITH project concludes my dissertation by demonstrating that network analysis is one way of disrupting existing paradigms for understanding the social-signification of ekphrastic poetry, but there are more methods available through computational tools such as text modeling, word frequency analysis, and classification that might also be useful.

To this end, Iâ€™ve begun by asking three modest questions about ekphrastic poetry using a machine learning application called MALLET:

1.) Could a computer learn to differentiate between ekphrastic poems by male and female poets?Â In â€œEkphrasis and the Other,â€ W.J.T. Mitchell argues that were we to read ekphrastic poems by women as opposed to ekphrastic poetry by men, that we might find a very different relationship between the active, speaking poetic voice and the passive, silent work of artâ€”a dynamic which informs our primary understanding of how ekphrastic poetry operates.Â Were this true and were the difference to occur within recurring topics and language use, a computer might be trained to recognize patterns more likely to co-occur in poetry by men or by women.

2.) Will topic modeling of ekphrastic texts pick out â€œstillnessâ€ as one of the most common topics in the genre?Â Much of the definition of ekphrasis revolves around the language of stillness: poetic texts, it has been argued, contemplate the stillness and muteness of the image with which it is engaged.Â Stillness, metaphorically linked to muteness, breathlessness, and death, provides one of the most powerful rationales for an understanding how words and images relate to one another within the ut pictura poesis traditionâ€”usually seen as an hostile encounter between rival forms of representation.Â The argument to this point has been made largely on critical interpretations enacted through close readings of a limited number of texts.Â Would a computer designed to recognize co-occurrences of words and assign those words to a â€œtopicâ€ based on the probability they would occur together also reveal a similar affiliation between stillness and death, muteness, even femininity?

3.) Would a computer be able to ascertain stylistic and semantic differences between ekphrastic and non-ekphrastic texts and reliably classify them according to whether or not the subject of the poem is an aesthetic object or not?Â We tend to believe that there are no real differences between how we describe the natural world as opposed to how we describe visual representations of the natural world.Â We base this assumption on human, interpretive, close readings ofÂ poetic texts; however, there is the potential that a computer might recognize subtle differences as statistically significant when considering hundreds of poems at a time.Â If a classification program such as Mallet could reliably categorize texts according to ekphrastic and non-ekphrastic, it is possible that we have missed something along the way.

In general, these are small questions constructed in such a way that there is a reasonable likelihood that we may get useful results.Â (I purposefully choose the word results instead of answers, because none of these would be answers.Â Instead the result of each study is designed to turn critics back to the texts with new questions.)Â And yet, how do we distinguish between useful results and something else?Â How do we know if it worked?Â Lots of money is spent trying to answer this question about big data, but what about these small and mid-sized data sets?Â Is there a threshold for how much data we need to be accurate and trustworthy?Â Can we actually develop standards for how much data we need to ask particular kinds of humanities questions to make relevant discoveries?Â In part, my project also addresses these questions, because otherwise, I canâ€™t make convincing arguments about the humanities questions Iâ€™m asking.

Small projects (even mid-sized projects with mid-sized datasets) offer the promise of richly encoded data that can be tested, reorganized, and applied flexibly to a variety of contexts without potentially becoming the entirety of a project directorâ€™s career.Â The space between close, highly-supervised readings and distant, unsupervised analysis remains wide open as a field of study, and yet its potential value as a manageable, not wholly consuming, and reproducible option make it worth seriously considering.Â What exactly can be accomplished by small and mid-scale projects is largely unknown, but it may well be that small and mid-sized projects are where many scholars will find the most satisfying and useful results.

Lisa @ Work

This site has moved as of March 20, 2013 to a new location: www.lisarhody.com

Category Archives: DH Tools

Why use visualizations to study poetry?

Chunks, Topics, and Themes in LDA

Small Projects & Limited Datasets