createdb iLund4uPlasmids/all_proteins.fa iLund4uPlasmids/mmseqs/DB/sequencesDB MMseqs Version: f6c98807d589091c625db68da258d587795acbab Database type 0 Shuffle input database true Createdb mode 0 Write lookup file 1 Offset of numeric ids 0 Compressed 0 Verbosity 3 Converting sequences [=================================================================================================== 1 Mio. sequences processed =================================================================================================== 2 Mio. sequences processed =================================================================================================== 3 Mio. sequences processed =================================================================================================== 4 Mio. sequences processed =================================================================================================== 5 Mio. sequences processed =================================================================================================== 6 Mio. sequences processed =================================================================================================== 7 Mio. sequences processed =================================================================================================== 8 Mio. sequences processed =================================================================================================== 9 Mio. sequences processed =================================================================================================== 10 Mio. sequences processed =================================================================================================== 11 Mio. sequences processed =================================================================================================== 12 Mio. sequences processed =================================================================================================== 13 Mio. sequences processed =========================================================================== Time for merging to sequencesDB_h: 0h 0m 2s 605ms Time for merging to sequencesDB: 0h 0m 4s 91ms Database type: Aminoacid Time for processing: 0h 0m 58s 17ms Create directory iLund4uPlasmids/mmseqs/DB/tmp cluster iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/clusterDB iLund4uPlasmids/mmseqs/DB/tmp --cluster-mode 0 --cov-mode 0 --min-seq-id 0.5 -c 0.75 -s 7 --max-seqs 41272320 MMseqs Version: f6c98807d589091c625db68da258d587795acbab Substitution matrix aa:blosum62.out,nucl:nucleotide.out Seed substitution matrix aa:VTML80.out,nucl:nucleotide.out Sensitivity 7 k-mer length 0 Target search mode 0 k-score seq:2147483647,prof:2147483647 Alphabet size aa:21,nucl:5 Max sequence length 65535 Max results per query 41272320 Split database 0 Split mode 2 Split memory limit 0 Coverage threshold 0.75 Coverage mode 0 Compositional bias 1 Compositional bias 1 Diagonal scoring true Exact k-mer matching 0 Mask residues 1 Mask residues probability 0.9 Mask lower case residues 0 Minimum diagonal score 15 Selected taxa Include identical seq. id. false Spaced k-mers 1 Preload mode 0 Pseudo count a substitution:1.100,context:1.400 Pseudo count b substitution:4.100,context:5.800 Spaced k-mer pattern Local temporary path Threads 48 Compressed 0 Verbosity 3 Add backtrace false Alignment mode 3 Alignment mode 0 Allow wrapped scoring false E-value threshold 0.001 Seq. id. threshold 0.5 Min alignment length 0 Seq. id. mode 0 Alternative alignments 0 Max reject 2147483647 Max accept 2147483647 Score bias 0 Realign hits false Realign score bias -0.2 Realign max seqs 2147483647 Correlation score weight 0 Gap open cost aa:11,nucl:5 Gap extension cost aa:1,nucl:2 Zdrop 40 Rescore mode 0 Remove hits by seq. id. and coverage false Sort results 0 Cluster mode 0 Max connected component depth 1000 Similarity type 2 Weight file name Cluster Weight threshold 0.9 Single step clustering false Cascaded clustering steps 3 Cluster reassign false Remove temporary files false Force restart with latest tmp false MPI runner k-mers per sequence 21 Scale k-mers per sequence aa:0.000,nucl:0.200 Adjust k-mer length false Shift hash 67 Include only extendable false Skip repeating k-mers false Set cluster iterations to 3 linclust iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --alph-size aa:13,nucl:5 --kmer-per-seq 21 --spaced-kmer-mode 1 --kmer-per-seq-scale aa:0.000,nucl:0.200 --adjust-kmer-len 0 --mask 0 --mask-prob 0.9 --mask-lower-case 0 -k 0 --hash-shift 67 --split-memory-limit 0 --include-only-extendable 0 --ignore-multi-kmer 0 --rescore-mode 0 --filter-hits 0 --sort-results 0 --remove-tmp-files 0 --force-reuse 0 kmermatcher iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --alph-size aa:13,nucl:5 --min-seq-id 0.5 --kmer-per-seq 21 --spaced-kmer-mode 1 --kmer-per-seq-scale aa:0.000,nucl:0.200 --adjust-kmer-len 0 --mask 0 --mask-prob 0.9 --mask-lower-case 0 --cov-mode 0 -k 0 -c 0.75 --max-seq-len 65535 --hash-shift 67 --split-memory-limit 0 --include-only-extendable 0 --ignore-multi-kmer 0 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 kmermatcher iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --alph-size aa:13,nucl:5 --min-seq-id 0.5 --kmer-per-seq 21 --spaced-kmer-mode 1 --kmer-per-seq-scale aa:0.000,nucl:0.200 --adjust-kmer-len 0 --mask 0 --mask-prob 0.9 --mask-lower-case 0 --cov-mode 0 -k 0 -c 0.75 --max-seq-len 65535 --hash-shift 67 --split-memory-limit 0 --include-only-extendable 0 --ignore-multi-kmer 0 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Database size: 13757440 type: Aminoacid Reduced amino acid alphabet: (A S T) (C) (D B N) (E Q Z) (F Y) (G) (H) (I V) (K R) (L J M) (P) (W) (X) Generate k-mers list for 1 split [=================================================================] 13.76M 24s 594ms Sort kmer 0h 0m 1s 491ms Sort by rep. sequence 0h 0m 0s 807ms Time for fill: 0h 0m 2s 60ms Time for merging to pref: 0h 0m 0s 0ms Time for processing: 0h 0m 32s 730ms rescorediagonal iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_rescore1 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --rescore-mode 0 --wrapped-scoring 0 --filter-hits 0 -e 0.001 -c 0.75 -a 0 --cov-mode 0 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --add-self-matches 0 --sort-results 0 --db-load-mode 0 --threads 48 --compressed 0 -v 3 [=================================================================] 13.76M 37s 171ms Time for merging to pref_rescore1: 0h 0m 5s 714ms Time for processing: 0h 0m 47s 983ms clust iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_rescore1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pre_clust --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 13.76M 3s 616ms Sort entries Find missing connections Found 33329436 new connections. Reconstruct initial order [=================================================================] 13.76M 3s 547ms Add missing connections [=================================================================] 13.76M 1s 449ms Time for read in: 0h 0m 10s 585ms Total time: 0h 0m 19s 21ms Size of the sequence database: 13757440 Size of the alignment database: 13757440 Number of clusters: 2928806 Writing results 0h 0m 0s 383ms Time for merging to pre_clust: 0h 0m 0s 0ms Time for processing: 0h 0m 20s 543ms createsubdb iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/order_redundancy iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/input_step_redundancy -v 3 --subdb-mode 1 Time for merging to input_step_redundancy: 0h 0m 0s 0ms Time for processing: 0h 0m 1s 282ms createsubdb iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/order_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_filter1 -v 3 --subdb-mode 1 Time for merging to pref_filter1: 0h 0m 0s 0ms Time for processing: 0h 0m 3s 725ms filterdb iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_filter1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_filter2 --filter-file iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/order_redundancy --threads 48 --compressed 0 -v 3 Filtering using file(s) [=================================================================] 2.93M 0s 781ms Time for merging to pref_filter2: 0h 0m 0s 988ms Time for processing: 0h 0m 2s 375ms rescorediagonal iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_filter2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_rescore2 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --rescore-mode 1 --wrapped-scoring 0 --filter-hits 1 -e 0.001 -c 0.75 -a 0 --cov-mode 0 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --add-self-matches 0 --sort-results 0 --db-load-mode 0 --threads 48 --compressed 0 -v 3 [=================================================================] 2.93M 5s 717ms Time for merging to pref_rescore2: 0h 0m 1s 259ms Time for processing: 0h 0m 8s 764ms align iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pref_rescore2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/aln --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 2928806 type: Aminoacid Target database size: 2928806 type: Aminoacid Calculation of alignments [=================================================================] 2.93M 13s 913ms Time for merging to aln: 0h 0m 1s 3ms 4071429 alignments calculated 3423488 sequence pairs passed the thresholds (0.840857 of overall calculated) 1.168902 hits per query sequence Time for processing: 0h 0m 16s 798ms clust iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/aln iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/clust --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 2.93M 0s 209ms Sort entries Find missing connections Found 494682 new connections. Reconstruct initial order [=================================================================] 2.93M 0s 207ms Add missing connections [=================================================================] 2.93M 0s 66ms Time for read in: 0h 0m 0s 755ms Total time: 0h 0m 1s 763ms Size of the sequence database: 2928806 Size of the alignment database: 2928806 Number of clusters: 2605698 Writing results 0h 0m 0s 206ms Time for merging to clust: 0h 0m 0s 0ms Time for processing: 0h 0m 2s 247ms mergeclusters iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/pre_clust iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/linclust/12612479980860989708/clust --threads 48 --compressed 0 -v 3 Clustering step 1 [=================================================================] 2.93M 1s 232ms Clustering step 2 [=================================================================] 2.61M 1s 774ms Write merged clustering [=================================================================] 13.76M 2s 476ms Time for merging to clu_redundancy: 0h 0m 1s 32ms Time for processing: 0h 0m 4s 465ms createsubdb iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_redundancy iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy -v 3 --subdb-mode 1 Time for merging to input_step_redundancy: 0h 0m 0s 0ms Time for processing: 0h 0m 1s 258ms prefilter iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/pref_step0 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --seed-sub-mat 'aa:VTML80.out,nucl:nucleotide.out' -s 1 -k 0 --target-search-mode 0 --k-score seq:2147483647,prof:2147483647 --alph-size aa:21,nucl:5 --max-seq-len 65535 --max-seqs 41272320 --split 0 --split-mode 2 --split-memory-limit 0 -c 0.75 --cov-mode 0 --comp-bias-corr 0 --comp-bias-corr-scale 1 --diag-score 0 --exact-kmer-matching 0 --mask 1 --mask-prob 0.9 --mask-lower-case 0 --min-ungapped-score 0 --add-self-matches 0 --spaced-kmer-mode 1 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --threads 48 --compressed 0 -v 3 Query database size: 2605698 type: Aminoacid Estimated memory consumption: 12G Target database size: 2605698 type: Aminoacid Index table k-mer threshold: 154 at k-mer size 6 Index table: counting k-mers [=================================================================] 2.61M 1s 916ms Index table: Masked residues: 6185102 Index table: fill [=================================================================] 2.61M 1s 855ms Index statistics Entries: 271533861 DB size: 2042 MB Avg k-mer size: 4.242717 Top 10 k-mers GPGGTL 5770 YATHQG 2687 YTGTPK 2648 EARGGR 1967 GPGGTT 1830 HQSGQR 1370 NVHTRW 1093 PHFGRQ 1007 ARRRGR 971 GQQVGR 947 Time for index table init: 0h 0m 4s 378ms Hard disk might not have enough free space (10T left).The prefilter result might need up to 259T. Process prefiltering step 1 of 1 k-mer similarity threshold: 154 Starting prefiltering scores calculation (step 1 of 1) Query db start 1 to 2605698 Target db start 1 to 2605698 [=================================================================] 2.61M 7s 822ms 1.764293 k-mers per position 3080 DB matches per sequence 0 overflows 58 sequences passed prefiltering per query sequence 17 median result list length 0 sequences with 0 size result lists Time for merging to pref_step0: 0h 0m 0s 972ms Time for processing: 0h 0m 19s 457ms align iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/pref_step0 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/aln_step0 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 0 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 2605698 type: Aminoacid Target database size: 2605698 type: Aminoacid Calculation of alignments [=================================================================] 2.61M 8m 17s 120ms Time for merging to aln_step0: 0h 0m 0s 931ms 71667970 alignments calculated 9305636 sequence pairs passed the thresholds (0.129844 of overall calculated) 3.571264 hits per query sequence Time for processing: 0h 8m 20s 21ms clust iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/aln_step0 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step0 --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 2.61M 0s 371ms Sort entries Find missing connections Found 5742 new connections. Reconstruct initial order [=================================================================] 2.61M 0s 408ms Add missing connections [=================================================================] 2.61M 0s 360ms Time for read in: 0h 0m 1s 474ms Total time: 0h 0m 4s 395ms Size of the sequence database: 2605698 Size of the alignment database: 2605698 Number of clusters: 2012877 Writing results 0h 0m 0s 171ms Time for merging to clu_step0: 0h 0m 0s 0ms Time for processing: 0h 0m 4s 876ms createsubdb iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step0 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 -v 3 --subdb-mode 1 Time for merging to input_step1: 0h 0m 0s 0ms Time for processing: 0h 0m 0s 433ms prefilter iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/pref_step1 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --seed-sub-mat 'aa:VTML80.out,nucl:nucleotide.out' -s 4 -k 0 --target-search-mode 0 --k-score seq:2147483647,prof:2147483647 --alph-size aa:21,nucl:5 --max-seq-len 65535 --max-seqs 41272320 --split 0 --split-mode 2 --split-memory-limit 0 -c 0.75 --cov-mode 0 --comp-bias-corr 1 --comp-bias-corr-scale 1 --diag-score 1 --exact-kmer-matching 0 --mask 1 --mask-prob 0.9 --mask-lower-case 0 --min-ungapped-score 15 --add-self-matches 0 --spaced-kmer-mode 1 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --threads 48 --compressed 0 -v 3 Query database size: 2012877 type: Aminoacid Estimated memory consumption: 9G Target database size: 2012877 type: Aminoacid Index table k-mer threshold: 127 at k-mer size 6 Index table: counting k-mers [=================================================================] 2.01M 1s 940ms Index table: Masked residues: 5058475 Index table: fill [=================================================================] 2.01M 2s 856ms Index statistics Entries: 439014609 DB size: 3000 MB Avg k-mer size: 6.859603 Top 10 k-mers GPGGTL 2583 YTGTPK 1838 AAALAG 1599 AAAAGA 1576 AAAAAR 1488 TSSGKG 1469 GEGGVV 1319 DDGGSL 1318 GEGGTL 1303 AAAAAP 1194 Time for index table init: 0h 0m 5s 550ms Hard disk might not have enough free space (10T left).The prefilter result might need up to 154T. Process prefiltering step 1 of 1 k-mer similarity threshold: 127 Starting prefiltering scores calculation (step 1 of 1) Query db start 1 to 2012877 Target db start 1 to 2012877 [=================================================================] 2.01M 2m 4s 300ms 54.303830 k-mers per position 77231 DB matches per sequence 22 overflows 548 sequences passed prefiltering per query sequence 358 median result list length 0 sequences with 0 size result lists Time for merging to pref_step1: 0h 0m 0s 753ms Time for processing: 0h 2m 17s 385ms align iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/pref_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/aln_step1 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 2012877 type: Aminoacid Target database size: 2012877 type: Aminoacid Calculation of alignments [=================================================================] 2.01M 8m 11s 360ms Time for merging to aln_step1: 0h 0m 0s 717ms 281441369 alignments calculated 2146056 sequence pairs passed the thresholds (0.007625 of overall calculated) 1.066164 hits per query sequence Time for processing: 0h 8m 15s 499ms clust iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/aln_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step1 --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 2.01M 0s 124ms Sort entries Find missing connections Found 16815 new connections. Reconstruct initial order [=================================================================] 2.01M 0s 122ms Add missing connections [=================================================================] 2.01M 0s 32ms Time for read in: 0h 0m 0s 460ms Total time: 0h 0m 3s 600ms Size of the sequence database: 2012877 Size of the alignment database: 2012877 Number of clusters: 1952397 Writing results 0h 0m 0s 154ms Time for merging to clu_step1: 0h 0m 0s 0ms Time for processing: 0h 0m 4s 39ms createsubdb iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step2 -v 3 --subdb-mode 1 Time for merging to input_step2: 0h 0m 0s 0ms Time for processing: 0h 0m 0s 390ms prefilter iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/pref_step2 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --seed-sub-mat 'aa:VTML80.out,nucl:nucleotide.out' -s 7 -k 0 --target-search-mode 0 --k-score seq:2147483647,prof:2147483647 --alph-size aa:21,nucl:5 --max-seq-len 65535 --max-seqs 41272320 --split 0 --split-mode 2 --split-memory-limit 0 -c 0.75 --cov-mode 0 --comp-bias-corr 1 --comp-bias-corr-scale 1 --diag-score 1 --exact-kmer-matching 0 --mask 1 --mask-prob 0.9 --mask-lower-case 0 --min-ungapped-score 15 --add-self-matches 0 --spaced-kmer-mode 1 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --threads 48 --compressed 0 -v 3 Query database size: 1952397 type: Aminoacid Estimated memory consumption: 9G Target database size: 1952397 type: Aminoacid Index table k-mer threshold: 100 at k-mer size 6 Index table: counting k-mers [=================================================================] 1.95M 1s 906ms Index table: Masked residues: 4944990 Index table: fill [=================================================================] 1.95M 2s 758ms Index statistics Entries: 434175746 DB size: 2972 MB Avg k-mer size: 6.783996 Top 10 k-mers AAAAAA 3486 AALAAA 2534 GPGGTL 2492 AAAAAL 2016 YTGTPK 1828 ALAALA 1661 AAALLA 1653 AAALAL 1612 ALAAAL 1529 AALAAL 1481 Time for index table init: 0h 0m 5s 408ms Hard disk might not have enough free space (10T left).The prefilter result might need up to 145T. Process prefiltering step 1 of 1 k-mer similarity threshold: 100 Starting prefiltering scores calculation (step 1 of 1) Query db start 1 to 1952397 Target db start 1 to 1952397 [=================================================================] 1.95M 36m 42s 236ms 956.969868 k-mers per position 1664266 DB matches per sequence 160421 overflows 21110 sequences passed prefiltering per query sequence 14286 median result list length 0 sequences with 0 size result lists Time for merging to pref_step2: 0h 0m 0s 935ms Time for processing: 0h 36m 54s 636ms align iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/pref_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/aln_step2 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 1952397 type: Aminoacid Target database size: 1952397 type: Aminoacid Calculation of alignments [=================================================================] 1.95M 1h 34m 23s 50ms Time for merging to aln_step2: 0h 0m 17s 39ms 8715499795 alignments calculated 1958699 sequence pairs passed the thresholds (0.000225 of overall calculated) 1.003228 hits per query sequence Time for processing: 1h 34m 46s 632ms clust iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/input_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/aln_step2 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step2 --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 1.95M 0s 93ms Sort entries Find missing connections Found 926 new connections. Reconstruct initial order [=================================================================] 1.95M 0s 97ms Add missing connections [=================================================================] 1.95M 0s 15ms Time for read in: 0h 0m 0s 387ms Total time: 0h 0m 1s 93ms Size of the sequence database: 1952397 Size of the alignment database: 1952397 Number of clusters: 1949021 Writing results 0h 0m 0s 154ms Time for merging to clu_step2: 0h 0m 0s 0ms Time for processing: 0h 0m 2s 369ms mergeclusters iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/clusterDB iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_redundancy iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step0 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step1 iLund4uPlasmids/mmseqs/DB/tmp/11626968526009775331/clu_step2 --threads 48 --compressed 0 -v 3 Clustering step 1 [=================================================================] 2.61M 11s 656ms Clustering step 2 [=================================================================] 2.01M 13s 712ms Clustering step 3 [=================================================================] 1.95M 15s 256ms Clustering step 4 [=================================================================] 1.95M 15s 752ms Write merged clustering [=================================================================] 13.76M 16s 436ms Time for merging to clusterDB: 0h 0m 2s 893ms Time for processing: 0h 0m 21s 742ms createtsv iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/sequencesDB iLund4uPlasmids/mmseqs/DB/clusterDB iLund4uPlasmids/mmseqs/mmseqs_clustering.tsv MMseqs Version: f6c98807d589091c625db68da258d587795acbab First sequence as representative false Target column 1 Add full header false Sequence source 0 Database output false Threads 48 Compressed 0 Verbosity 3 Time for merging to mmseqs_clustering.tsv: 0h 0m 12s 599ms Time for processing: 0h 0m 40s 621ms