July 13 2007

Astrophysics Dataset: Pandora

Inizio i lavori sul dataset Pandora fornitomi dal prof. Longo, basandomi sulle sue direttive.

  • Verrà usato un sottoinsieme delle colonne
  • Un primo clustering verrà effettuato depurando il dataset da missing values
  • Un successivo clustering verrà effettuato sul dataset non depurato
  • I due clustering verranno confrontati, utilizzando il primo come baseline di riferimento.
  • Maggiori dettagli sul dataset saranno disponibili al più presto.

    Il clustering verrà affrontato con Bregman Co-clustering, per affrontare il problema dei missing values.
    Il metodo di aggiornamento dei mediodi/centroidi sarà il Local Search, che evita minimi locali e ci permette, partendo da un numero iniziale sovrastimato di cluter, di “scovare” il numero effettivo di cluster (o nei casi difficili una buona approssimazione di esso), lavorando per raffinamenti successivi.
    In questo esperimento l’inizializzazione del co-clustering sarà lasciata casuale.

    In successive prove proveremo ad utilizzare l’inizializzazione spettrale proposta in

    • H. Cho, I. Dhillon, Y. Guan, and S. Sra, "Minimum sum squared residue co-clustering of gene expression data," in Proceedings of the Fourth SIAM International Conference on Data Mining, 2004, pp. 114-125.
      @inproceedings{cho04minimum,
        author = {H. Cho and I. Dhillon and Y. Guan and S. Sra},
        Booktitle = {Proceedings of the Fourth SIAM International Conference on Data Mining},
        Date-Added = {2007-04-12 11:30:35 +0200},
        Date-Modified = {2007-06-19 15:14:55 +0200},
        Keywords = {clustering, co-clustering, bioinformatics},
        Month = {April},
        Pages = {114–125},
        Title = {Minimum sum squared residue co-clustering of gene expression data},
        Url = {http://www.cs.utexas.edu/users/inderjit/public_papers/mssrcc_siam.pdf},
        Year = {2004},
        Bdsk-File-1 = {YnBsaXN0MDDUAQIDBAUGBwpZJGFyY2hpdmVyWCR2ZXJzaW9uVCR0b3BYJG9iamVjdHNfEA9OU0tleWVkQXJjaGl2ZXISAAGGoNEICVRyb290gAGoCwwXGBkaHiVVJG51bGzTDQ4PEBMWWk5TLm9iamVjdHNXTlMua2V5c1YkY2xhc3OiERKABIAFohQVgAKAA4AHXHJlbGF0aXZlUGF0aFlhbGlhc0RhdGFfEFkuLi8uLi8uLi9QYXBlcnMvQ2hvL01pbmltdW0gc3VtIHNxdWFyZWQgcmVzaWR1ZSBjby1jbHVzdGVyaW5nIG9mIGdlbmUgZXhwcmVzc2lvbiBkYXRhLnBkZtIbDxwdV05TLmRhdGFPEQJYAAAAAAJYAAIAAAlEb2N1bWVudHMAAAAAAAAAAAAAAAAAAAAAAAC+zniuSCsAAAA3JQAfTWluaW11bSBzdW0gc3F1YXJlZCAjMkEzOTY0LnBkZgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACo5ZMI6vZZQREYgAAAAAAADAAMAAAkAAAAAAAAAAAAAAAAAAAAAA0NobwAAEAAIAAC+zlyOAAAAEQAIAADCOqF2AAAAAQAUADclAAA3G4AAALLyAAASxgAAEq0AAgBORG9jdW1lbnRzOm5lbW86RG9jdW1lbnRzOlVuaXZlcnNpdGE6UGFwZXJzOkNobzpNaW5pbXVtIHN1bSBzcXVhcmVkICMyQTM5NjQucGRmAA4AjABFAE0AaQBuAGkAbQB1AG0AIABzAHUAbQAgAHMAcQB1AGEAcgBlAGQAIAByAGUAcwBpAGQAdQBlACAAYwBvAC0AYwBsAHUAcwB0AGUAcgBpAG4AZwAgAG8AZgAgAGcAZQBuAGUAIABlAHgAcAByAGUAcwBzAGkAbwBuACAAZABhAHQAYQAuAHAAZABmAA8AFAAJAEQAbwBjAHUAbQBlAG4AdABzABIAay9uZW1vL0RvY3VtZW50cy9Vbml2ZXJzaXRhL1BhcGVycy9DaG8vTWluaW11bSBzdW0gc3F1YXJlZCByZXNpZHVlIGNvLWNsdXN0ZXJpbmcgb2YgZ2VuZSBleHByZXNzaW9uIGRhdGEucGRmAAATABIvVm9sdW1lcy9Eb2N1bWVudHMAFQACABf//wAAgAbSHyAhIlgkY2xhc3Nlc1okY2xhc3NuYW1loyIjJF1OU011dGFibGVEYXRhVk5TRGF0YVhOU09iamVjdNIfICYnoickXE5TRGljdGlvbmFyeQAIABEAGwAkACkAMgBEAEkATABRAFMAXABiAGkAdAB8AIMAhgCIAIoAjQCPAJEAkwCgAKoBBgELARMDbwNxA3YDfwOKA44DnAOjA6wDsQO0AAAAAAAAAgEAAAAAAAAAKAAAAAAAAAAAAAAAAAAAA8E=},
        Bdsk-Url-1 = {http://www.cs.utexas.edu/users/inderjit/public_papers/mssrcc_siam.pdf}
      }

    per migliorare la qualità del risultato finale.

    Infine, essendo presenti valori negativi nella matrice, l’istanza di Co-clustering basata su di divergenza KL e Mutua Informazione non potrà essere utilizzata

    • I. S. Dhillon, S. Mallela, and D. S. Modha, "Information-Theoretic Co-Clustering," in Proceedings of The Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2003), 2003, pp. 89-98.
      @inproceedings{dhillon:mallela:modha:03,
        author = {I. S. Dhillon and S. Mallela and D. S. Modha},
        Booktitle = {Proceedings of The Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining ({KDD}-2003)},
        Date-Modified = {2007-07-14 15:32:35 +0200},
        Keywords = {clustering, co-clustering, relative entropy},
        Pages = {89–98},
        Title = {Information-Theoretic Co-Clustering},
        Url = {http://www.cs.utexas.edu/users/inderjit/public_papers/kdd_cocluster.pdf},
        Year = {2003},
        Bdsk-File-1 = {YnBsaXN0MDDUAQIDBAUGBwpZJGFyY2hpdmVyWCR2ZXJzaW9uVCR0b3BYJG9iamVjdHNfEA9OU0tleWVkQXJjaGl2ZXISAAGGoNEICVRyb290gAGoCwwXGBkaHiVVJG51bGzTDQ4PEBMWWk5TLm9iamVjdHNXTlMua2V5c1YkY2xhc3OiERKABIAFohQVgAKAA4AHXHJlbGF0aXZlUGF0aFlhbGlhc0RhdGFfED8uLi8uLi8uLi9QYXBlcnMvRGhpbGxvbi9JbmZvcm1hdGlvbi1UaGVvcmV0aWMgQ28tQ2×1c3RlcmluZy5wZGbSGw8cHVdOUy5kYXRhTxECCgAAAAACCgACAAAJRG9jdW1lbnRzAAAAAAAAAAAAAAAAAAAAAAAAvs54rkgrAAAANyNdH0luZm9ybWF0aW9uLVRoZW9yZXRpIzIzQThBNi5wZGYAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAjqKbCBy5VAAAAAAAAAAAAAwADAAAJAAAAAAAAAAAAAAAAAAAAAAdEaGlsbG9uAAAQAAgAAL7OXI4AAAARAAgAAMIHIEUAAAABABQANyNdADcbgAAAsvIAABLGAAASrQACAFJEb2N1bWVudHM6bmVtbzpEb2N1bWVudHM6VW5pdmVyc2l0YTpQYXBlcnM6RGhpbGxvbjpJbmZvcm1hdGlvbi1UaGVvcmV0aSMyM0E4QTYucGRmAA4AUAAnAEkAbgBmAG8AcgBtAGEAdABpAG8AbgAtAFQAaABlAG8AcgBlAHQAaQBjACAAQwBvAC0AQwBsAHUAcwB0AGUAcgBpAG4AZwAuAHAAZABmAA8AFAAJAEQAbwBjAHUAbQBlAG4AdABzABIAUS9uZW1vL0RvY3VtZW50cy9Vbml2ZXJzaXRhL1BhcGVycy9EaGlsbG9uL0luZm9ybWF0aW9uLVRoZW9yZXRpYyBDby1DbHVzdGVyaW5nLnBkZgAAEwASL1ZvbHVtZXMvRG9jdW1lbnRzABUAAgAX//8AAIAG0h8gISJYJGNsYXNzZXNaJGNsYXNzbmFtZaMiIyRdTlNNdXRhYmxlRGF0YVZOU0RhdGFYTlNPYmplY3TSHyAmJ6InJFxOU0RpY3Rpb25hcnkACAARABsAJAApADIARABJAEwAUQBTAFwAYgBpAHQAfACDAIYAiACKAI0AjwCRAJMAoACqAOwA8QD5AwcDCQMOAxcDIgMmAzQDOwNEA0kDTAAAAAAAAAIBAAAAAAAAACgAAAAAAAAAAAAAAAAAAANZ},
        Bdsk-Url-1 = {http://www.cs.utexas.edu/users/inderjit/public_papers/kdd_cocluster.pdf}
      }

    Post a comment

    This blog is multi language by p.osting.it's Babel