5. Web Tools
ZcurveHub provides various tools for analyzing genomes using Z-curve method, which are freely available on https://tubic.tju.edu.cn/zcurve/.
5.1. Z-curve Plotter
Z-curve Plotter allows users to upload 1 to 3 nucleotide sequences and select 1 to 4 components from 11 components for visualization for each sequence at most. With Plotly’s powerful interactive charts, you can scale, rotate curve graphs and view information about each coordinate point. Auxiliary curves (AT/GC skew/fraction) is also supported for comparative analysis.
5.1.1. Input
We offer 3 input methods: File Upload, Text Input and NCBI Accession. Each input method should meet the above requirements.
5.1.1.1. File Upload
FASTA and GenBank files are allowed to be uploaded.
Note: The annotations in the GenBank file will be automatically ignored and it only retains the nucleic acid sequence information.

5.1.1.2. FASTA Text Input
Upload the nucleic acid sequence in FASTA text format.
Note: Illegal characters in the sequence will be regarded as “N” and count it into the total length of the sequence.

5.1.1.3. NCBI Accession
Enter a legal NCBI Nucleotide accession number (e.g. KP205272.1), and the program can automatically download the sequence from the public database.

5.1.2. Settings
Z-curve Plotter manages the input sequences in a card-like manner and enables customized settings for all items conveniently.
5.1.2.1. 2D Mode
The card displays basic information such as the name, length, and GC content of the sequence, as well as options such as the plot start point, plot end point, smoothing window size, and curve types. If your input provides a matching NCBI accession, you can view the Sequence features using the NCBI Sequence Viewer.
Start
Plot start point. Its index starting from 0 like most programming language, and should not be larger than the length of the sequence.End
Plot end point. Its value requirement is the same as that of Start.Window
The window size of mean smoothing. The range of values is [0, 1000].2D Curve Types
Optional curve types, up to 4 types can be selected for each sequence. The definition of each type can be found here.Extra Options
Set parameters for the auxiliary curves (AT/GC skew/fraction).

5.1.2.2. 3D Mode
The information displayed on this card is basically the same as that in the 2D mode, but the option of the curve type is limited to the three coordinate axes of the 3D chart.
3D Curve Types
Any 3 of the 11 components can be selected as the X-axis component, Y-axis component and Z-axis component, but do not choose the same components as different axes, because the result would be trivial.

5.1.3. Visualization
The visualized results can be saved as PNG images via Plotly, or the original data can be saved as JSON through the download button below.
5.1.3.1. 2D Mode

5.1.3.2. 3D Mode

5.2. Z-curve Encoder
5.2.1. Input
Z-curve Encoder allows users to upload sequences or genomes in FASTA/GenBank format and can choose to use the sequences directly as input or extract CDSs from GenBank as input. Our official server requires that the length of a single sequence cannot exceed 30 kb, and the total number of sequences cannot exceed 50,000.
5.2.1.1. Sequence File Upload
The default input for this tool is not allowed to be empty or to submit an empty file. When the machine learning option is turned on, this entry will default to positive sample input. The result of this feature extraction is in [job_id]_features.npy in the downloaded output file.

5.2.1.2. Negative Control Upload
Additional input options for this application. If machine learning is enabled, this option is treated as a negative example control data set, but if machine learning is not enabled, it is treated as just another data set. The result of this feature extraction is in [job_id]_negctrl_features.npy in the downloaded output file.

5.2.1.3. Negative Control Shuffling
An option to generate a set of random sequences. For the ZCURVE system, it is a means of generating negative examples during self-training, while for the machine learning process, it can verify the validity of the model. The result of this feature extraction is in [job_id]_shuffle_features.npy in the downloaded output file.
Negative/Positive Ratio
The ratio between the generated negative and positive samples.Random Seed
A seed used to generate random numbers.

5.2.2. Settings
Sets the hyperparameters for the phase-dependent K-nucleotide Z-curve transformation. This module allows the user to freely combine Z-curve parameters (up to 6 layers) to find the most suitable feature extraction scheme. This table can automatically calculate the number of parameters generated by each layer transformation, so that users can design deep learning models.
K-nucleotide
The length of the K-nucleotides to be counted and used in Z-curve transformation (k = 1, 2, 3, 4, 5, 6).Phase
Total phase number of phase specific Z-curve transformation (phase =1, 2, 3, 4, 5, 6).Frequencized
Whether to enable the frequencized Z-curve parameter. The default mode is the global frequency method.
\(p_i({\rm N}_{k-1}{\rm X})=n_i({\rm N}_{k-1}{\rm X})/(n - k + 1)\)Local
Use the local frequency method.
\(p_i({\rm N}_{k-1}{\rm X})=n_i({\rm N}_{k-1}{\rm X})/(n_i({\rm N}_{k-1}{\rm A})+n_i({\rm N}_{k-1}{\rm G})+n_i({\rm N}_{k-1}{\rm C})+n_i({\rm N}_{k-1}{\rm T}))\)
Add A New Row to increase the number of layers, move the key Up and Down to change the order in which the parameters appear, and the key Delete to reduce the number of layers.

5.2.3. Preprocessing
Data preprocessing options, available using the API provided by Sci-kit Learn. When checked, the output will overwrite the original encoding result and output the corresponding sklearn object. For example, if the standardization option is checked, an additional [job_id]_std_scaler.joblib will appear in the output.
Standardization
To extract the features of standardizing operations, will callsklearn.preprocessing.StandardScalerNormalization
To extract the features of normalizing operations, will callsklearn.preprocessing.MinMaxScalerPrincipal Component Analysis (PCA)
Perform principal component analysis for dimensionality reduction on the data.Components
The number of components retained in the data after dimensionality reduction.
K-means Clustering
Conduct unsupervised cluster analysis on the original data or preprocessed data.Clusters
Specify the number of clusters.

5.2.4. Machine Learning
Here we provide three models that learn the Z-curve parameters’ features well: Support Vector Machine (SVM), Random Forest (RF) and Multilayer Perceptron (MLP). Each of them provides two hyperparameters that have the most significant impact on the model effect for adjustment. In addition, the weights of positive and negative samples can also be set freely. The final machine learning models will be saved as [job_id]_ml_model.joblib and can be used like sci-kit learn models.

5.2.5. Output
The task of Z-curve Encoder is computationally intensive, so it cannot feed back the results in real time like the other two tools. You need to get the results from the Job ID generated when you submit it. In the job details page, you can get the basic information when submitting.

5.2.5.1. Visualization
Here we provide visualizations of principal component analysis and cluster analysis (PCA must be checked previously, otherwise cluster analysis will only output the metrics). You can intuitively see the power of the Z-curve method for sequence feature extraction.

5.2.5.2. Metrics
We provide five machine learning metrics and four clustering metrics for reference. The AUC metric provides a visual representation.
5.3. Z-curve Segmenter
5.3.1. Input
Z-curve Segmenter allows users to enter in three ways: plain text, file, and NCBI Accession. Either way, only one sequence can be entered. For example, if a user submits a file with multiple sequences, only the first sequence will be taken, and the rest, along with annotation information, will be ignored. Our official server requires that the length of a single sequence cannot exceed 100 kb.
5.3.1.1. Text Input
The text input field can highlight the A, T, G, and C characters to help you determine the correctness of the input sequence, but only spaces, line feeds, tabs, and carriage returns are filtered out when submitted. Characters such as R, Y, M, K, W, S, B, D, H, V, U are processed normally, and other characters are treated as N (including N).

5.3.1.2. File Input
Both FASTA and GenBank inputs are supported, but only one file can be uploaded. If the file contains multiple nucleic acid sequences, the back-end program will only process the first one. All annotations from GenBank will be ignored.

5.3.1.3. Accession Input
Our program supports the online acquisition of sequences from NCBI, but please be careful to enter the correct Accession, not the protein sequence or assembly.

5.3.2. Options
Select and set the mode of segmentation and the parameters of the algorithm. The Z-curve segmentation method was based on genome order index, and it is a binary iterative algorithm. The segment points in each iteration are calculated by the following formula:
\(n_{\rm seg}={\rm argmax}\{w_1S({\rm P}_n) + w_2S({\rm Q}_n) - S(w_1{\rm P}_n + w_2{\rm Q}_n)\},n=1,2,3...,N\)
In the above formula, P represents the left subsequence at point \(N\) of a dna sequence of length \(n\), and Q represents the right subsequence.
Mode
We provide 7 ways to calculate S:Segmentation Target
Order Index S(P)
Application
Z-curve
\(S({\rm P})=a^2+g^2+c^2+t^2\)
Replication Origin Recognition
RY disparity
\(S({\rm P})=(a^2+g^2)+(c^2+t^2)\)
Mitochondrial rRNA Region Search
MK disparity
\(S({\rm P})=(a^2+c^2)+(g^2+t^2)\)
Mitochondrial \(\rm O_L\) Recognition
WS disparity
\(S({\rm P})=(a^2+t^2)+(g^2+c^2)\)
Genomic Island Search
AT disparity
\(S({\rm P})=a^2+t^2\)
GC disparity
\(S({\rm P})=g^2+c^2\)
Leading/Lagging Chain search
CpG profile
\(S({\rm P})=[p_n({\rm CpG})]^2+[1-p_n({\rm CpG})]^2\)
CpG Island Search
Note: ‘GC profile’ mode is the same as ‘WS disparity’, but it will output the negative form of the curve when visualized.
Start Position
The start position of the fragment [start, stop) to be segmented.
(The subscript starts at 0, as in most computer languages)End Position
The stop position of the fragment [start, stop) to be segmented.
(The subscript starts at 0, as in most computer languages)Note
If start is greater than end, the program defaults the topological nature of the sequence to a ring and rotates the sequence to the new starting point.Smoothing Window
The window size used for mean smoothing can reduce the graph sawtooth and make it more beautiful。Halting Parameter
In a round of iteration, if the \(n_{seg}\) value of the target fragment is less than this value, the iteration stops.Max Iterations
Maximum iterations. Prevents users from setting too small a halting parameter and getting stuck in infinite iterations.Min Length
The minimum distance between two segment points. Stop iteration when the distance is less than this.

5.3.3. Visualizaion
The results of the corresponding curve segmentation are visualized, so that users can adjust the parameters. With the powerful Plotly library, you can move, zoom, and hide curves to get the best picture possible. Sequence features can also be viewed if a valid NCBI accession is provided by the input.

5.3.4. Results
Display segment results in tabular form. The table is interactive, so you can click directly to jump to the corresponding region of the sequence. You can download it via the Download button. Use the Plotly control to save the image display.
