Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Monday, April 3, 2017

Language Modelling Datasets and Tools

An introduction to language modelling (LM) is here
For RNN LMs see Mikolov's slides here
Mikolov's RNN LM tool kit  is here
SRI international's language modelling tool kit is here

Language modelling can be slow with RNNs. This is a faster implementation that uses Eigen.

There is a 1B token benchmark dataset released by Google for evaluating language models. It can be obtained here It has tokenised and splitted heldout and train portions. Shard 0000 is used for reporting test results.


Monday, September 22, 2014

Distributed vs. Distributional Representations

Distributed Representations = Learnt using Deep Learning methods. Distributed over the parameter space. Dense, lower dimensional embeddings.

Distributional Representations = Counted from a corpus. Classical distributional hypothesis. Sparse, High dimensional.

Saturday, September 20, 2014

My referencing workflow

I use BibDesk for reference management. After a major academic event such as a conference, I access the proceedings website using BibDesk. ACL anthology indexes all papers for NLP conferences, IEEE Explorer can be used for IEEE conferences, ACM digital library works for ACM conferences, and Google Scholar in general for any. The BibTex entries can be imported to BibDesk from within BibDesk with a url to the paper. That is it. I then assign papers a To READ tag (keyword) or to a Prority Reading list. I then read those papers and assign keywords and annotations. I use Skim for reading and annotating PDFs on Mac.

For PDFs I find online, I send them to CiteULike and then import them to BibDesk from there. CiteULike can extract most of the attributes from PDFs (scientific papers) such as title and authors. I can use BibDesk again to access CiteULike and import these references.

Mendely is also a good tool if you want to find related papers.

This is my workflow.

Thursday, February 25, 2010

Interesting WWW 2010 papers

  • EXPLORING WEB SCALE LANGUAGE MODELS FOR SEARCH QUERY PROCESSING
    Jian Huang, Jiangbo Miao, Xiaolong Li, Jianfeng Gao and Kuansan Wang
  • CROSS-DOMAIN SENTIMENT CLASSIFICATION VIA SPECTRAL FEATURE ALIGNMENT
    Sinno Pan, Xiaochuan Ni, Jian-Tao Sun, Qiang Yang and Zheng Chen
  • BUILDING TAXONOMY OF WEB SEARCH INTENTS FOR NAME ENTITY QUERIES
    Xiaoxin Yin and Sarthak Shah
  • MULTI-MODALITY IN ONE-CLASS CLASSIFICATION
    Boris Chidlovskii and Matthijs Hovelynck
  • A SCALABLE MACHINE LEARNING APPROACH FOR SEMI-STRUCTURED NAMED ENTITY RECOGNITION
    Utku Irmak and Reiner Kraft
  • FACETED EXPLORATION OF IMAGE SEARCH RESULTS
    Roelof van Zwol and Börkur Sigurbjörnsson
  • TOWARDS NATURAL QUESTION GUIDED SEARCH
    Alexander Kotov and ChengXiang Zhai
  • A LARGE SCALE ACTIVE LEARNING SYSTEM FOR TOPICAL CATEGORIZATION ON THE WEB
    Suju Rajan, Dragomir Yankov, Scott Gaffney and Adwait Ratnaparkhi
  • DIVERSIFYING WEB SEARCH RESULTS
    Davood Rafiei, Krishna Bharat and Anand Shukla
  • A GENERAL FRAMEWORK FOR EXPLORING CATEGORY INFORMATION FOR QUESTION RETRIEVAL IN COMMUNITY QUESTION ANSWER ARCHIVES
    Xin Cao, Gao Cong, Bin Cui and Christian Jensen
  • RANKING SPECIALIZATION FOR WEB SEARCH: A DIVIDE-AND-CONQUER APPROACH BY USING TOPICAL RANKSVM
    Jiang Bian, Xin Li, Fan Li, Zhaohui Zheng and Hongyuan Zha
  • GENERALIZED DISTANCES BETWEEN RANKINGS
    Ravi Kumar and Sergei Vassilvitskii
  • THE ANATOMY OF A LARGE-SCALE SOCIAL SEARCH ENGINE
    Damon Horowitz and Sepandar Kamvar
  • USE TWITTER DATA FOR RECENCY RANKING IMPROVEMENT IN WEB SEARCH
    Anlei Dong, Ruiqiang Zhang, Pranam Kolari, Bai Jing, Yi Chang, Fernando Diaz, Zhaohui Zheng and Hongyuan Zha
  • A CHARACTERIZATION OF ONLINE SEARCH BEHAVIOR
    Ravi Kumar and Andrew Tomkins
  • WHAT IS TWITTER, A SOCIAL NETWORK OR A NEWS MEDIA?
    Haewoon Kwak, Changhyun Lee, Hosung Park and Sue Moon

Saturday, September 26, 2009

classias notes

There is an excellent software to train/predict a range of ML algorithms including logistic regression with L1 or L2 regularization, pagasos SVM L1/L2, average perceptron. It is called Classias and is by Naoaki Okazaki.
It is amazingly fast and can handle large datasets! It work directly on compressed formats such as bz2, tar.gz

Here are some quick how to note
  • running binary logistic regression 

    classias-train -tb -a lbfgs.logistic -m

    -tb says it is of -t type binary b.
    -a specifies the learning algorithm, which is lbfgs optimized logistic regression (note: logistic regression is NOT a regression model but a classification algorithm), current L1/L2 regularization is supported only with lbfgs.
    -m  specifies the model file.
    the final entry is the actual training file. The format being,
    label fid:fval ...

  • cross validation, regularization and help
    To perform cross validation use -g5 -x options (5 says 5-fold cross validation, can use any integer there.) If you have your held out data on a separate file then you can specify both training and heldout data files using another set of options [see the documentation of Classias]. If you want to enable L1 regularization use -pc1=1 This says set the parameter c1 to the value 1 (regularization coefficient) for the algorithm specified by -a. The default value is zero for L1 regularization and 1 for L2. If you are using L1 regularization only then you must set L2 to zero. i.e. -pc2=0. Otherwise you will end up using both L1 and L2 regularizations! For example, if you set both regularization coefficients to 1, then you end up having more features in the final trained model compared to what you get if trained only with L1 regularization. But still, it is far less features than what you would get if you used only L2 regularization. For the RCV1 dataset, I got 40628 features only usng L2 (accuracy being 0.95), where as those values were 491(@0.95) for L1 only and 1597(@0.94) using both.

    General help of classias-train can be seen by doing,
    classias-train --help
    and to see what parameters are available for a specific algorithm (e.g. lbfgs.logistic) do the following,
    classias-train -a lbfgs.logistic -H
    H indicates parameter specific help. -h is the normal help.
    Putting it all together the following command trains a binary logistic regression model with L1 regularization and also performs 5-fold cross validation.
     classias-train -tb -a lbfgs.logistic -pc1=1 -pc2=0  -m rcv1.binary.model -g5 -x rcv1_train.binary

  • Multi-class classification
  • Tagging  (prediction)
    Read test instances from stdin and output the class labels , weights (-w), and in the case of logistic regression models probabilities (-p). Specify the model file by -m. You can compute accuracies by using -t option. To suppress labels etc. when testing use quiet option (-q).

    cat rcv1_test_binary | classias-tag -m rcv1.binary.model -p cat rcv1_test_binary | classias-tag -m rcv1.binary.model -w cat rcv1_test_binary | classias-tag -m rcv1.binary.model -tq
    If your data is in bz2 the use bzcat instead of cat.  

Sunday, June 21, 2009

Useful Links

http://www-nlp.stanford.edu/~mgalley/
Lexical Chains tool is available here.

Convert source code to pure HTML. Useful when pasting code on a blog.
http://pygments.org/

Continuously monitor GPU usage

 For nvidia GPUs do the follwing: nvidia-smi -l 1