createdb iLund4uPhages/all_proteins.fa iLund4uPhages/mmseqs/DB/sequencesDB MMseqs Version: f6c98807d589091c625db68da258d587795acbab Database type 0 Shuffle input database true Createdb mode 0 Write lookup file 1 Offset of numeric ids 0 Compressed 0 Verbosity 3 Converting sequences [=================================================================================================== 1 Mio. sequences processed =================================================================================================== 2 Mio. sequences processed =================================================================================================== 3 Mio. sequences processed =================================================================================================== 4 Mio. sequences processed =================================================================================================== 5 Mio. sequences processed =================================================================================================== 6 Mio. sequences processed =================================================================================================== 7 Mio. sequences processed =================================================================================================== 8 Mio. sequences processed =================================================================================================== 9 Mio. sequences processed =================================================================================================== 10 Mio. sequences processed =================================================================================================== 11 Mio. sequences processed =================================================================================================== 12 Mio. sequences processed =================================================================================================== 13 Mio. sequences processed =================================================================================================== 14 Mio. sequences processed =================================================================================================== 15 Mio. sequences processed =================================================================================================== 16 Mio. sequences processed =================================================================================================== 17 Mio. sequences processed =================================================================================================== 18 Mio. sequences processed =================================================================================================== 19 Mio. sequences processed =================================================================================================== 20 Mio. sequences processed =================================================================================================== 21 Mio. sequences processed =================================================================================================== 22 Mio. sequences processed =================================================================================================== 23 Mio. sequences processed =================================================================================================== 24 Mio. sequences processed =================================================================================================== 25 Mio. sequences processed =================================================================================================== 26 Mio. sequences processed =================================================================================================== 27 Mio. sequences processed =================================================================================================== 28 Mio. sequences processed =================================================================================================== 29 Mio. sequences processed =================================================================================================== 30 Mio. sequences processed =================================================================================================== 31 Mio. sequences processed =================================================================================================== 32 Mio. sequences processed =================================================================================================== 33 Mio. sequences processed =================================================================================================== 34 Mio. sequences processed =================================================================================================== 35 Mio. sequences processed =================================================================================================== 36 Mio. sequences processed =================================================================================================== 37 Mio. sequences processed =================================================================================================== 38 Mio. sequences processed =================================================================================================== 39 Mio. sequences processed =================================================================================================== 40 Mio. sequences processed =================================================================================================== 41 Mio. sequences processed =================================================================================================== 42 Mio. sequences processed =================================================== Time for merging to sequencesDB_h: 0h 0m 7s 816ms Time for merging to sequencesDB: 0h 0m 12s 928ms Database type: Aminoacid Time for processing: 0h 2m 29s 802ms Create directory iLund4uPhages/mmseqs/DB/tmp cluster iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/clusterDB iLund4uPhages/mmseqs/DB/tmp --cluster-mode 0 --cov-mode 0 --min-seq-id 0.5 -c 0.75 -s 7 --max-seqs 127544253 MMseqs Version: f6c98807d589091c625db68da258d587795acbab Substitution matrix aa:blosum62.out,nucl:nucleotide.out Seed substitution matrix aa:VTML80.out,nucl:nucleotide.out Sensitivity 7 k-mer length 0 Target search mode 0 k-score seq:2147483647,prof:2147483647 Alphabet size aa:21,nucl:5 Max sequence length 65535 Max results per query 127544253 Split database 0 Split mode 2 Split memory limit 0 Coverage threshold 0.75 Coverage mode 0 Compositional bias 1 Compositional bias 1 Diagonal scoring true Exact k-mer matching 0 Mask residues 1 Mask residues probability 0.9 Mask lower case residues 0 Minimum diagonal score 15 Selected taxa Include identical seq. id. false Spaced k-mers 1 Preload mode 0 Pseudo count a substitution:1.100,context:1.400 Pseudo count b substitution:4.100,context:5.800 Spaced k-mer pattern Local temporary path Threads 48 Compressed 0 Verbosity 3 Add backtrace false Alignment mode 3 Alignment mode 0 Allow wrapped scoring false E-value threshold 0.001 Seq. id. threshold 0.5 Min alignment length 0 Seq. id. mode 0 Alternative alignments 0 Max reject 2147483647 Max accept 2147483647 Score bias 0 Realign hits false Realign score bias -0.2 Realign max seqs 2147483647 Correlation score weight 0 Gap open cost aa:11,nucl:5 Gap extension cost aa:1,nucl:2 Zdrop 40 Rescore mode 0 Remove hits by seq. id. and coverage false Sort results 0 Cluster mode 0 Max connected component depth 1000 Similarity type 2 Weight file name Cluster Weight threshold 0.9 Single step clustering false Cascaded clustering steps 3 Cluster reassign false Remove temporary files false Force restart with latest tmp false MPI runner k-mers per sequence 21 Scale k-mers per sequence aa:0.000,nucl:0.200 Adjust k-mer length false Shift hash 67 Include only extendable false Skip repeating k-mers false Set cluster iterations to 3 linclust iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --alph-size aa:13,nucl:5 --kmer-per-seq 21 --spaced-kmer-mode 1 --kmer-per-seq-scale aa:0.000,nucl:0.200 --adjust-kmer-len 0 --mask 0 --mask-prob 0.9 --mask-lower-case 0 -k 0 --hash-shift 67 --split-memory-limit 0 --include-only-extendable 0 --ignore-multi-kmer 0 --rescore-mode 0 --filter-hits 0 --sort-results 0 --remove-tmp-files 0 --force-reuse 0 kmermatcher iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --alph-size aa:13,nucl:5 --min-seq-id 0.5 --kmer-per-seq 21 --spaced-kmer-mode 1 --kmer-per-seq-scale aa:0.000,nucl:0.200 --adjust-kmer-len 0 --mask 0 --mask-prob 0.9 --mask-lower-case 0 --cov-mode 0 -k 0 -c 0.75 --max-seq-len 65535 --hash-shift 67 --split-memory-limit 0 --include-only-extendable 0 --ignore-multi-kmer 0 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 kmermatcher iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --alph-size aa:13,nucl:5 --min-seq-id 0.5 --kmer-per-seq 21 --spaced-kmer-mode 1 --kmer-per-seq-scale aa:0.000,nucl:0.200 --adjust-kmer-len 0 --mask 0 --mask-prob 0.9 --mask-lower-case 0 --cov-mode 0 -k 0 -c 0.75 --max-seq-len 65535 --hash-shift 67 --split-memory-limit 0 --include-only-extendable 0 --ignore-multi-kmer 0 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Database size: 42514394 type: Aminoacid Reduced amino acid alphabet: (A S T) (C) (D B N) (E Q Z) (F Y) (G) (H) (I V) (K R) (L J M) (P) (W) (X) Generate k-mers list for 1 split [=================================================================] 42.51M 59s 204ms Sort kmer 0h 0m 5s 10ms Sort by rep. sequence 0h 0m 2s 495ms Time for fill: 0h 0m 5s 971ms Time for merging to pref: 0h 0m 0s 1ms Time for processing: 0h 1m 23s 769ms rescorediagonal iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_rescore1 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --rescore-mode 0 --wrapped-scoring 0 --filter-hits 0 -e 0.001 -c 0.75 -a 0 --cov-mode 0 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --add-self-matches 0 --sort-results 0 --db-load-mode 0 --threads 48 --compressed 0 -v 3 [=================================================================] 42.51M 1m 42s 823ms Time for merging to pref_rescore1: 0h 0m 15s 131ms Time for processing: 0h 2m 9s 855ms clust iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_rescore1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pre_clust --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 42.51M 28s 271ms Sort entries Find missing connections Found 120758933 new connections. Reconstruct initial order [=================================================================] 42.51M 27s 133ms Add missing connections [=================================================================] 42.51M 6s 32ms Time for read in: 0h 1m 8s 162ms Total time: 0h 1m 24s 16ms Size of the sequence database: 42514394 Size of the alignment database: 42514394 Number of clusters: 5190738 Writing results 0h 0m 0s 862ms Time for merging to pre_clust: 0h 0m 0s 0ms Time for processing: 0h 1m 28s 383ms createsubdb iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/order_redundancy iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/input_step_redundancy -v 3 --subdb-mode 1 Time for merging to input_step_redundancy: 0h 0m 0s 0ms Time for processing: 0h 0m 3s 539ms createsubdb iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/order_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_filter1 -v 3 --subdb-mode 1 Time for merging to pref_filter1: 0h 0m 0s 0ms Time for processing: 0h 0m 7s 625ms filterdb iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_filter1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_filter2 --filter-file iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/order_redundancy --threads 48 --compressed 0 -v 3 Filtering using file(s) [=================================================================] 5.19M 1s 871ms Time for merging to pref_filter2: 0h 0m 2s 166ms Time for processing: 0h 0m 5s 217ms rescorediagonal iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_filter2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_rescore2 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --rescore-mode 1 --wrapped-scoring 0 --filter-hits 1 -e 0.001 -c 0.75 -a 0 --cov-mode 0 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --add-self-matches 0 --sort-results 0 --db-load-mode 0 --threads 48 --compressed 0 -v 3 [=================================================================] 5.19M 7s 834ms Time for merging to pref_rescore2: 0h 0m 2s 75ms Time for processing: 0h 0m 13s 865ms align iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pref_rescore2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/aln --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 5190738 type: Aminoacid Target database size: 5190738 type: Aminoacid Calculation of alignments [=================================================================] 5.19M 39s 979ms Time for merging to aln: 0h 0m 1s 947ms 7568688 alignments calculated 6109048 sequence pairs passed the thresholds (0.807148 of overall calculated) 1.176913 hits per query sequence Time for processing: 0h 0m 46s 22ms clust iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/aln iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/clust --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 5.19M 0s 643ms Sort entries Find missing connections Found 918310 new connections. Reconstruct initial order [=================================================================] 5.19M 0s 547ms Add missing connections [=================================================================] 5.19M 0s 168ms Time for read in: 0h 0m 1s 879ms Total time: 0h 0m 3s 962ms Size of the sequence database: 5190738 Size of the alignment database: 5190738 Number of clusters: 4587328 Writing results 0h 0m 0s 384ms Time for merging to clust: 0h 0m 0s 0ms Time for processing: 0h 0m 4s 906ms mergeclusters iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/pre_clust iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/linclust/4856118789553293960/clust --threads 48 --compressed 0 -v 3 Clustering step 1 [=================================================================] 5.19M 4s 816ms Clustering step 2 [=================================================================] 4.59M 5s 720ms Write merged clustering [=================================================================] 42.51M 6s 875ms Time for merging to clu_redundancy: 0h 0m 1s 705ms Time for processing: 0h 0m 11s 519ms createsubdb iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_redundancy iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy -v 3 --subdb-mode 1 Time for merging to input_step_redundancy: 0h 0m 0s 0ms Time for processing: 0h 0m 3s 600ms prefilter iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/pref_step0 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --seed-sub-mat 'aa:VTML80.out,nucl:nucleotide.out' -s 1 -k 0 --target-search-mode 0 --k-score seq:2147483647,prof:2147483647 --alph-size aa:21,nucl:5 --max-seq-len 65535 --max-seqs 127544253 --split 0 --split-mode 2 --split-memory-limit 0 -c 0.75 --cov-mode 0 --comp-bias-corr 0 --comp-bias-corr-scale 1 --diag-score 0 --exact-kmer-matching 0 --mask 1 --mask-prob 0.9 --mask-lower-case 0 --min-ungapped-score 0 --add-self-matches 0 --spaced-kmer-mode 1 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --threads 48 --compressed 0 -v 3 Query database size: 4587328 type: Aminoacid Estimated memory consumption: 20G Target database size: 4587328 type: Aminoacid Index table k-mer threshold: 154 at k-mer size 6 Index table: counting k-mers [=================================================================] 4.59M 2s 977ms Index table: Masked residues: 17136214 Index table: fill [=================================================================] 4.59M 3s 49ms Index statistics Entries: 447170249 DB size: 3047 MB Avg k-mer size: 6.987035 Top 10 k-mers ECRQPD 2336 GNGGSP 2197 LDFGTT 1991 CEPDSY 1584 NKLCFQ 1469 IGTHGK 1456 HIAKEN 1453 FSYPSR 1447 FIRIIW 1440 PQIREW 1419 Time for index table init: 0h 0m 6s 881ms Hard disk might not have enough free space (10T left).The prefilter result might need up to 803T. Process prefiltering step 1 of 1 k-mer similarity threshold: 154 Starting prefiltering scores calculation (step 1 of 1) Query db start 1 to 4587328 Target db start 1 to 4587328 [=================================================================] 4.59M 13s 944ms 1.973140 k-mers per position 3984 DB matches per sequence 0 overflows 66 sequences passed prefiltering per query sequence 21 median result list length 0 sequences with 0 size result lists Time for merging to pref_step0: 0h 0m 1s 755ms Time for processing: 0h 0m 34s 16ms align iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/pref_step0 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/aln_step0 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 0 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 4587328 type: Aminoacid Target database size: 4587328 type: Aminoacid Calculation of alignments [=================================================================] 4.59M 23m 16s 855ms Time for merging to aln_step0: 0h 0m 1s 849ms 131153716 alignments calculated 21046694 sequence pairs passed the thresholds (0.160473 of overall calculated) 4.588007 hits per query sequence Time for processing: 0h 23m 23s 249ms clust iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/aln_step0 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step0 --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 4.59M 0s 928ms Sort entries Find missing connections Found 41196 new connections. Reconstruct initial order [=================================================================] 4.59M 0s 824ms Add missing connections [=================================================================] 4.59M 2s 58ms Time for read in: 0h 0m 4s 494ms Total time: 0h 0m 8s 334ms Size of the sequence database: 4587328 Size of the alignment database: 4587328 Number of clusters: 3517610 Writing results 0h 0m 0s 317ms Time for merging to clu_step0: 0h 0m 0s 0ms Time for processing: 0h 0m 9s 9ms createsubdb iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step0 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 -v 3 --subdb-mode 1 Time for merging to input_step1: 0h 0m 0s 0ms Time for processing: 0h 0m 0s 786ms prefilter iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/pref_step1 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --seed-sub-mat 'aa:VTML80.out,nucl:nucleotide.out' -s 4 -k 0 --target-search-mode 0 --k-score seq:2147483647,prof:2147483647 --alph-size aa:21,nucl:5 --max-seq-len 65535 --max-seqs 127544253 --split 0 --split-mode 2 --split-memory-limit 0 -c 0.75 --cov-mode 0 --comp-bias-corr 1 --comp-bias-corr-scale 1 --diag-score 1 --exact-kmer-matching 0 --mask 1 --mask-prob 0.9 --mask-lower-case 0 --min-ungapped-score 15 --add-self-matches 0 --spaced-kmer-mode 1 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --threads 48 --compressed 0 -v 3 Query database size: 3517610 type: Aminoacid Estimated memory consumption: 15G Target database size: 3517610 type: Aminoacid Index table k-mer threshold: 127 at k-mer size 6 Index table: counting k-mers [=================================================================] 3.52M 2s 913ms Index table: Masked residues: 12480377 Index table: fill [=================================================================] 3.52M 4s 384ms Index statistics Entries: 707510272 DB size: 4536 MB Avg k-mer size: 11.054848 Top 10 k-mers GNGGSP 2064 RISGGI 1602 NLWTEN 1469 TLEALG 1291 NHDNNN 1215 GGGGGG 1015 IDSNIG 962 AAALAG 939 LDFGTT 935 NVITPS 928 Time for index table init: 0h 0m 8s 250ms Hard disk might not have enough free space (10T left).The prefilter result might need up to 472T. Process prefiltering step 1 of 1 k-mer similarity threshold: 127 Starting prefiltering scores calculation (step 1 of 1) Query db start 1 to 3517610 Target db start 1 to 3517610 [=================================================================] 3.52M 3m 55s 356ms 60.737111 k-mers per position 112765 DB matches per sequence 0 overflows 804 sequences passed prefiltering per query sequence 541 median result list length 0 sequences with 0 size result lists Time for merging to pref_step1: 0h 0m 1s 337ms Time for processing: 0h 4m 15s 514ms align iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/pref_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/aln_step1 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 3517610 type: Aminoacid Target database size: 3517610 type: Aminoacid Calculation of alignments [=================================================================] 3.52M 14m 28s 454ms Time for merging to aln_step1: 0h 0m 1s 201ms 597069378 alignments calculated 3866366 sequence pairs passed the thresholds (0.006476 of overall calculated) 1.099146 hits per query sequence Time for processing: 0h 14m 35s 290ms clust iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/aln_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step1 --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 3.52M 0s 338ms Sort entries Find missing connections Found 39722 new connections. Reconstruct initial order [=================================================================] 3.52M 0s 351ms Add missing connections [=================================================================] 3.52M 0s 61ms Time for read in: 0h 0m 1s 140ms Total time: 0h 0m 5s 27ms Size of the sequence database: 3517610 Size of the alignment database: 3517610 Number of clusters: 3368967 Writing results 0h 0m 0s 316ms Time for merging to clu_step1: 0h 0m 0s 0ms Time for processing: 0h 0m 5s 704ms createsubdb iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step2 -v 3 --subdb-mode 1 Time for merging to input_step2: 0h 0m 0s 0ms Time for processing: 0h 0m 0s 660ms prefilter iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/pref_step2 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' --seed-sub-mat 'aa:VTML80.out,nucl:nucleotide.out' -s 7 -k 0 --target-search-mode 0 --k-score seq:2147483647,prof:2147483647 --alph-size aa:21,nucl:5 --max-seq-len 65535 --max-seqs 127544253 --split 0 --split-mode 2 --split-memory-limit 0 -c 0.75 --cov-mode 0 --comp-bias-corr 1 --comp-bias-corr-scale 1 --diag-score 1 --exact-kmer-matching 0 --mask 1 --mask-prob 0.9 --mask-lower-case 0 --min-ungapped-score 15 --add-self-matches 0 --spaced-kmer-mode 1 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --threads 48 --compressed 0 -v 3 Query database size: 3368967 type: Aminoacid Estimated memory consumption: 15G Target database size: 3368967 type: Aminoacid Index table k-mer threshold: 100 at k-mer size 6 Index table: counting k-mers [=================================================================] 3.37M 2s 845ms Index table: Masked residues: 12187359 Index table: fill [=================================================================] 3.37M 4s 465ms Index statistics Entries: 696118832 DB size: 4471 MB Avg k-mer size: 10.876857 Top 10 k-mers AAAAAA 2890 GNGGSP 2061 AAVAAA 1506 NLWTEN 1430 GAGGDD 1366 AAAAAL 1307 NHDNNN 1182 AAAAAV 1057 LAAGLA 1012 GGGGGG 990 Time for index table init: 0h 0m 8s 272ms Hard disk might not have enough free space (10T left).The prefilter result might need up to 433T. Process prefiltering step 1 of 1 k-mer similarity threshold: 100 Starting prefiltering scores calculation (step 1 of 1) Query db start 1 to 3368967 Target db start 1 to 3368967 [=================================================================] 3.37M 1h 11m 12s 747ms 1041.560225 k-mers per position 2479835 DB matches per sequence 224045 overflows 31374 sequences passed prefiltering per query sequence 19961 median result list length 0 sequences with 0 size result lists Time for merging to pref_step2: 0h 0m 1s 481ms Time for processing: 1h 11m 32s 686ms align iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/pref_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/aln_step2 --sub-mat 'aa:blosum62.out,nucl:nucleotide.out' -a 0 --alignment-mode 3 --alignment-output-mode 0 --wrapped-scoring 0 -e 0.001 --min-seq-id 0.5 --min-aln-len 0 --seq-id-mode 0 --alt-ali 0 -c 0.75 --cov-mode 0 --max-seq-len 65535 --comp-bias-corr 1 --comp-bias-corr-scale 1 --max-rejected 2147483647 --max-accept 2147483647 --add-self-matches 0 --db-load-mode 0 --pca substitution:1.100,context:1.400 --pcb substitution:4.100,context:5.800 --score-bias 0 --realign 0 --realign-score-bias -0.2 --realign-max-seqs 2147483647 --corr-score-weight 0 --gap-open aa:11,nucl:5 --gap-extend aa:1,nucl:2 --zdrop 40 --threads 48 --compressed 0 -v 3 Compute score, coverage and sequence identity Query database size: 3368967 type: Aminoacid Target database size: 3368967 type: Aminoacid Calculation of alignments [=================================================================] 1.00M 47m 16s 241ms [=================================================================] 1.00M 47m 34s 910ms [=================================================================] 1.00M 47m 45s 788ms [=================================================================] 368.97K 17m 29s 131ms Time for merging to aln_step2: 0h 0m 1s 199ms 18301295278 alignments calculated 3389857 sequence pairs passed the thresholds (0.000185 of overall calculated) 1.006201 hits per query sequence Time for processing: 2h 40m 20s 50ms clust iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/input_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/aln_step2 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step2 --cluster-mode 0 --max-iterations 1000 --similarity-type 2 --threads 48 --compressed 0 -v 3 --cluster-weight-threshold 0.9 Clustering mode: Set Cover [=================================================================] 3.37M 0s 251ms Sort entries Find missing connections Found 3258 new connections. Reconstruct initial order [=================================================================] 3.37M 0s 261ms Add missing connections [=================================================================] 3.37M 0s 30ms Time for read in: 0h 0m 1s 143ms Total time: 0h 0m 2s 422ms Size of the sequence database: 3368967 Size of the alignment database: 3368967 Number of clusters: 3357888 Writing results 0h 0m 0s 276ms Time for merging to clu_step2: 0h 0m 0s 0ms Time for processing: 0h 0m 3s 142ms mergeclusters iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/clusterDB iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_redundancy iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step0 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step1 iLund4uPhages/mmseqs/DB/tmp/9232704515845691129/clu_step2 --threads 48 --compressed 0 -v 3 Clustering step 1 [=================================================================] 4.59M 7s 82ms Clustering step 2 [=================================================================] 3.52M 8s 112ms Clustering step 3 [=================================================================] 3.37M 9s 13ms Clustering step 4 [=================================================================] 3.36M 9s 886ms Write merged clustering [=================================================================] 42.51M 11s 989ms Time for merging to clusterDB: 0h 0m 1s 132ms Time for processing: 0h 0m 18s 281ms createtsv iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/sequencesDB iLund4uPhages/mmseqs/DB/clusterDB iLund4uPhages/mmseqs/mmseqs_clustering.tsv MMseqs Version: f6c98807d589091c625db68da258d587795acbab First sequence as representative false Target column 1 Add full header false Sequence source 0 Database output false Threads 48 Compressed 0 Verbosity 3 Time for merging to mmseqs_clustering.tsv: 0h 0m 4s 951ms Time for processing: 0h 0m 14s 248ms