Look into clustering and clustering summarization #225

ctb · 2017-05-16T17:02:21Z

things to provide code for --

notebooks to investigate & viz
include relabling (by editing labels.txt / lists in Python)
cut clusters at some distance (cophenetic?)
bootstrap/summarize/shrink clusters
histogram of distances
selection of samples by fishing with a query (or query set)

ctb · 2017-05-17T12:59:42Z

@ekg suggested looking into variational autoencoding:

Or, if you're interested in finding your way back into the big tree, you could use VAE or similar dim. > reduction and work from that repr

(we already have t-SNE working in a notebook somewhere)

ctb · 2021-03-04T14:35:00Z

see cluster and cocluster too #700

@mr-eyes

This PR adds a new command, `cluster`, that can be used to cluster the output from `pairwise` and `multisearch`. `cluster`uses `rustworkx-core` (which internally uses `petgraph`) to build a graph, adding edges between nodes when the similarity exceeds the user-defined threshold. It can work on any of the similarity columns output by `pairwise` or `multisearch`, and will add all nodes to the graph to preserve singleton 'clusters' in the output. `cluster` outputs two files: 1. cluster identities file: `Component_X, name1;name2;name3...` 2. cluster size histogram `cluster_size, count` context for some things I tried: - try using petgraph directly and removing rustworkx dependency > nope,`rustworkx-core` adds `connected_components` that returns the connected components, rather than just the number of connected components. Could reimplement if `rustworkx-core` brings in a lot of deps - try using 'extend_with_edges' instead of add_edge logic. > nope, only in `petgraph` **Punted Issues:** - develop clustering visualizations (ref @mr-eyes kSpider/dbretina work). Optionally output dot file of graph? (#248) - enable updating clusters, rather than always regenerating from scratch (#249) - benchmark `cluster` (#247) > `pairwise` files can be millions of lines long. Would it be faster to parallel read them, store them in an `edges` vector, and then add nodes/edges sequentially? Note that we would probably need to either 1. store all edges, including those that do not pass threshold) or 2. After building the graph from edges, add nodes from `names_to_node` that are not already in the graph to preserve singletons. Related issues: * #219 * sourmash-bio/sourmash#2271 * sourmash-bio/sourmash#700 * sourmash-bio/sourmash#225 * sourmash-bio/sourmash#274 --------- Co-authored-by: C. Titus Brown <titus@idyll.org>

ctb mentioned this issue Jun 5, 2017

use alternative clustering package in sourmash plot, to support larger data sets #274

Open

ctb mentioned this issue Feb 26, 2024

MRG: Add graph-based clustering sourmash-bio/sourmash_plugin_branchwater#234

Merged

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Look into clustering and clustering summarization #225

Look into clustering and clustering summarization #225

ctb commented May 16, 2017

ctb commented May 17, 2017

ctb commented Mar 4, 2021

Look into clustering and clustering summarization #225

Look into clustering and clustering summarization #225

Comments

ctb commented May 16, 2017

ctb commented May 17, 2017

ctb commented Mar 4, 2021