2011. október 5.

Könyvismertető: Data Mining with Rattle and R

Egyre többen érdeklődnek a az adattudományi (data science) és gépi tanulási módszerek iránt. Az adatbányászat napjainkban nem annyira felkapott téma, ahogy sokan igyekeznek elkerülni a statisztika és számítógépes statisztika (computatuional statistics) kifejezéseket, de megnyugtatunk mindenkit, a sok buzzword tkp. ugyanazt a dolgot fedi. A megnövekedett érdeklődés és a tény hogy életünket egyre jobban átszövik az említett területek eredményei együtt járnak az igénnyel egy egyszerű, gyakorlatorientált bevezetőre. Williams könyve remekül használható akár a programozásban kevésbé jártas, a statisztika alapjait ismerő érdeklődőknek.

2011. október 4.

An Introduction to Scientific Workflows (For Linguists)

A guest post by Richard Littauer

If you've been following @richlitt on Twitter for the past six months, you may have noticed that I've been talking about scientific workflows a lot. I was doing an internship for DataONE, an NSF-funded cyberinfrastructure initiative that tasked me with finding out all I could about scientific workflows from a site called myExperiment, which is a repository and social network for scientists who use them. However, if you're outside of the fields of bioinformatics or harder sciences, you may not know what I'm talking about when I talk about 'Scientific Workflows'.

2011. október 2.

Sublexical Semantics

A guest post by Richard Littauer

I was asked by Zoltan a long time ago to write something for this post, and have so far neglected to come up with anything. I'm about to start a two year Computational Linguistics masters at the University of Saarbrücken, so I figure it is about time I do this, before I am too bogged down to do anything. So, here is some original research I did a couple of years ago for a Lexical Semantics assignment at the University of Edinburgh, while I was in my undergraduate, being taught by Nikolas Gisbourne. His research in this area is largely within the framework of Word Grammar, and he specialises in the event structure of perception verbs (which is the title of his 2010 book released with OUP. I had wanted to take his word grammar and apply it some sort of binary system, but my thoughts quickly went in a different direction. I cover a lot of ground responding to Rappaport Hovav and Levin's work in this are, as well as Pustejovsky, mostly as I didn't want to risk quoting Gisbourne wrong at the time. Here is that different direction, then - it's very rough, and I had no computational background at the time, so it might be a bit out there. Hopefully, I'll get some feedback on this, though, as I think it's interesting and might be an interesting route to pursue. (NB: It's mostly edited from a longer essay, if it sound a bit formal.)

2011. október 1.

Adatbázis építés gyorsan és egyszerűen, függetlenül

Vajon hol húzódik az a határ, amikor egy alkalmazásunk elengedhetetlenül rászorul arra, hogy adatainkat külön, egy külső adatbázisba tároljuk? - A határ meghúzásához nem közvetlen választ szeretnék adni. Helyette a leírásban arra törekszem, hogy megmutassam, hogy hogyan lehet gyorsan, egyszerűen, adatbázis típusától független megoldásokat készíteni. Ezzel serkentve, és közvetve válaszolni a kérdésre: ahol lehetséges, használjuk bátran az adatbázisok által kínált megoldásokat. Válasszuk szét a logikát, a megvalósítást és felhasznált adatokat.