2012. április 7.
Fonémák és hasonulások?
2011. augusztus 10.
Nyelvi értelmező, házilag
2011. augusztus 7.
Könyvismertető: Számítógéppel emberi nyelven
- Prószéky Gábor – Kis Balázs: Számítógéppel emberi nyelven. Intelligens szövegkezelés számítógéppel
- Szak Kiadó, Budapest, 1999
2011. augusztus 4.
Könyvismertető: Az Üveghegyen innen
- Csernicskó Istávn - Kontra Miklós (szerk.): Az üveghegyen innen - Anyanyelvváltozatok, identitás és magyar anyanyelvi nevelés
- PoliPrint, II. Rákóczi Ferenc KMF, 2008
- Elektronikus verzió a Magyar Elektronikus könyvtár oldalán: http://mek.oszk.hu/08100/08144/cimkes.html
2011. április 5.
Hogyan kezdtem szófajelemzőt írni?
2010. november 12.
Bohumil Hrabal szótára
2010. augusztus 9.
The Anxiety of Digital Humanities
Digital humanities is an anxiety-ridden set of practices at the intersection of humanities research and computer technology. But the worst thing that could befall DH is forced collective psychotherapy or free prescriptions for Prozac. As long as we are anxious, we will try to find new and interesting things to do.
Zoltan asked me to write a post about my work in field of digital humanities (DH) and I am happy to do so. Not because I know ahead of time what I will say, but because thinking about one's own work -- and by thinking, i mean: reasoning more or less comprehensibly, without necessarily writing a multi-volume, self-aggrandizing novel -- is a fun and useful exercise. Mostly for myself.
DH is an anxiety-ridden set of practices at the intersection of humanities research and computer technology. It is an exciting but troubled discipline-in-the-making, uncertain about its own boundaries and purpose. Is it a field or a fad? Does it have a future or is it the future? Is it "emerging" or simply "peripheral"? Practitioners of DH are no weirder than your average academic (who is, by the standards of the "outside world" already pretty weird), but I think that an average DHer hears more voices in their head than a typical Slavist, for instance. With some notable exceptions.
The DH community is obsessed with trying to justify what it is that they are doing. A great deal of that soul-searching is unfortunately not very soulful: it is prompted by academic power games, grant opportunities and vicious self-promotion. But some aspects of this disciplinary (and sometimes undisciplined) introspection are fascinating: What should we do with a million books? Is there such a thing as a philosophy of text encoding? What are the limits of digital representation? What is digital history?
Academia in general is anxiety central: a place where insecure and often snobbish people disguise their distaste for manual labor as a kind of intellectual and quasi-moral superiority. The anxiety of digital humanities, however, is better: it is the anxiety about the very basics of professional intellectual work and models of representation and self-representation. Not being entirely certain about one's status or level of academic acceptance keeps one alert, active and safely tucked away from the oceanic feeling of complacency. That is why the worst thing that could happen to DH is forced collective psychotherapy or free prescriptions for Prozac. As long as we are anxious, we will try to find new and interesting things to do.
In my own work, I focus on a traditional tool of humanistic research: the dictionary, both as a material object, cultural product and a model of language. In lexicographic literature, you will find very little anxiety about what a dictionary is: it is usually defined as a list of words with some kind of explanation attached to them. For me, however, the dictionary is first and foremost a kind of text. As such, it is already a problem: a meaning potential that can be realized though its use, but also a field of contradictions that can not always be reconciled. That is why my work is focused on the interplay between electronic textuality and our notion of what a dictionary is (and ought to be). I am exploring ways in which the methods of digital humanities and digital libraries could alter our idea of what a dictionary can (and should) do.
At the Belgrade Center for Digital Humanities, we are working on a Wordnet-based bilingualized Serbian-English dictionary that will be deployed as a web service to interact with digital libraries. That project is called Transpoetika. We have also started digitalizing Serbian historical dictionaries with the goal of exploring the creation of a Serbian meta-dictionary. I am interested not only in the interaction of digital texts and digital dictionaries, but also in lexicographic serendipity: the Transpoetika dictionary, for instance, uses Twitter feeds as sources of "live quotes" and tagged Flickr images as on-the-fly illustration. Our experiment in lexicographic community-building called Reklakaza.la ("Hearsay") has drawn more than 22,000 fans on Facebook.
And, still, we are only at the very beginning. I have no idea how far we will be able to go and where we will end up. I don't know which of our experiments will be successful and which will fail: lexicographic text mining or dictionary visualizations? Spatial mappings of the lexicon or ludic explorations of the dictionary's narrative potential? What I do know for certain is that I have become good friends with my own anxiety and that I plan to keep that relationship going as long as I can.
Digital humanities should embrace their anxiety, too.
About the author
Toma Tasovac has a B.A. in Russian Literature from Harvard and M.A. in Comparative Literature from Princeton. He is the director of the Belgrade Center for Digital Humanities, a media trainer for DW-Akademie in Berlin and Bon, and a self-proclaimed dictionary freak. He meddles into all sorts of things, mostly digital.

2010. július 11.
NooJ, az Integrált Nyelvelemző Környezet I.
VENDÉGPOSZT!
Előzetesnek szánom ezt a cikket. Bemutatni, hogy mennyi és milyen minőségű nyelvelemző programok állnak már jelenleg is a rendelkezésünkre. De előzetes abban az értelemben is, hogy az eszközökről szeretnék majd több hosszabb-rövidebb leírást is adni. És nem utolsó sorban előzetes abból a szempontból is, hogy a későbbre tervezett Natural Language Toolkit használatát fogjuk a most bemutatott NooJ eszközzel megalapozni.
A NooJ, és elődje az INTEX egy integrált nyelvelemző környezet. Egy francia nyelvész készítette, aki ráébredt arra, hogy rengeteg szakterület tudná alkalmazni, használni a saját céljaira egy nyelvelemző rendszert. A többi nyelvelemző ellen mindig az első ellenérv a kezelhetőség volt. Még a jelenleg a PTE-n fejlesztett (prolog nyelven fejlesztett) szövegelemzőről is (bár csak néhány bemutatót sikerült erről a fejlesztésről megszereznem...) első hátrányként említik, hogy nehezen kezelhető, olvasható az eredmény. Ez azért könnyedén orvosolható lenne. Mindenesetre megértem azokat, akik csak egy-egy ötletért nem hajlandóak ennyire belemerülni a témában. Pont a számukra lehet a legideálisabb eszköz a NooJ.
Szerencsére a témának van magyar honlapja. http://corpus.nytud.hu/nooj/ címen tudjátok elérni. Itt található meg hozzá továbbá a magyar modul is, amivel el tudjuk végezni az elemzéseinket. Illetve a kipróbáláshoz ajánlom mindenki figyelmébe Vajda Péter bemutatását: http://corpus.nytud.hu/manye/vp_nooj.ppt
Hogy van magyar honlapja, ez sajnos nem egyenlő azzal, hogy fejlesztik is. 2006-ban indult, és azóta csak a magyar modul került fel. De a még akkoriban tervezett grammatika nem jelent meg azóta sem.
A NooJ a morphdb.hu-t használja. Akik használták már külön, azoknak nem lesz meglepetés a szavak elemzésének eredménye vagy a típushibái, de ezeket könnyedén javíthatjuk az aktuális szövegnél. Viszont a NooJ nem csak erre képes. Lehetséges vele szógyakoriságot vizsgálni, ahogyan lehet csak simán szegmentálni. Saját nyelvtannal kiegészítve pedig határ a csillagos ég. És akkor még nem is beszültünk arról, hogy képes az elemzett szöveget xml-formátumban visszaadni, tehát az eredményen tovább dolgozhatunk például az NLTK-val... de erről majd csak később.
Egy-két javaslatot azért tennék a program használatához. A Huntoken mondatszegmentálót használja. Ezért a legjobb eredmény érdekében minden sorban csak egyetlen mondat szerepeljen! Továbbá mivel a szavak elemzéséhez a Hunmorph-ot használja, így nem számítsunk eredményre a tulajdonnevek és a szóösszetételek esetén. Ezeket nem tudja kezelni.
Továbbá álljon itt egy minta is. Példaként és a várható eredmények előrejelzése végett. Ezt a cikket elemeztettem le vele, egészen eddig a bekezdésig:
Összesen 26 mondat
242 különböző szóalak
30 olyan szó, aminek nem tudta meghatározni a szófajtát, felépítését (tulajdonnevek, formátumtípusok, webcímek és szóösszetételek)
426 különböző felismert és elemzett szóalak (a kettő szám azért nem egyezik, az összes szóalak és az elemzett szóalakok száma, mert sok olyan szóalak van, ahol elképzelhető több elemzési eredmény a szövegkörnyezetnek megfelelően, de a program jelzi számunkra a lehetséges elemzéseket. Például a „szánom” két elemzési módja a következő: szánom,szán (szótő): N+nom+1+sg+pssg+ps vagy V+1+def+sg. Itt az emberi értelem meg tudja határozni, hogy az igei a helyes, de ezt csak jelentéstanilag tehetjük meg. Egy másik szövegkörnyezetben már főnévként szerepelhet.)
Szerző:
Gerő Dávid: Magyar és nyelvtechnológus hallgató, kezdő programozó és webfejlesztő, aki érdeklődik a nyelvészet és az informatika iránt. A határterületekért különösen rajongok, de sajnos mindkettőben csak kezdő, érdeklődő laikus vagyok.

