# Billing and Subscription Metrics -------------------------------------------------------------------------------- title: "Introduction" description: "Create a powerful ML pipeline that analyzed billing and subscription metrics for an OTT platform using the Catalyst QuickML components." last_updated: "2026-09-29T06:07:16.300Z" source: "https://docs.catalyst.zoho.com/en/tutorials/billing-and-subscription-metrics/introduction/" service: "All Services" -------------------------------------------------------------------------------- # OTT Platform: Billing and Subscription Metrics Tutorial ## Introduction This tutorial will guide you through building a time series Forecasting model using **Catalyst QuickML** to predict future billing and subscription metrics. Using daily historical data from the OTT platform’s billing system, the model aims to forecast trends such as total revenue, new subscriptions, cancellations, upgrades, and coupon usage. We will provide you with a sample dataset that can be used as the data source for the model. Before proceeding with the steps, let’s understand the basics of time series forecasting. Time series refers to a sequence of data points representing how one or more features change over time. These values are recorded at regular intervals, which allows for the analysis of trends, seasonal patterns, and anomalies within the data. Time series data can be univariate, where only a single feature (such as daily logins) is tracked, or multivariate, where multiple features (such as sessions, screen views, and file uploads) are recorded together over time. The primary objective of this analysis is to predict future billing and subscription outcomes. This helps businesses make informed decisions about revenue planning, churn management, promotional effectiveness, and capacity forecasting. By identifying temporal dependencies and usage patterns, the model provides actionable insights for optimizing subscription strategies and overall financial performance. To learn more about time series models, components of time series, and evaluation metrics, refer to the **help document**. Now, let's have a quick overview of the tutorial. ## Overview **1. Preprocess the dataset** The dataset contains daily billing and subscription metrics from the OTT platform, including features such as Total Revenue, New Subscriptions, Cancelled Subscriptions, Trial Conversions, Refunds Issued, Upgrades, and Coupon Usage. For this tutorial, we’ll use these multiple variables to build a multivariate forecasting model that captures the interdependencies between different business metrics over time. In Catalyst QuickML, the first step is to configure preprocessing. The Date column is designated as the timestamp, ensuring that the model correctly interprets the chronological sequence of data points. The dataset is aligned to a daily frequency to maintain consistent time intervals. Because the dataset contains missing values, imputation is applied to fill gaps in the data and preserve temporal continuity. For data normalization, a Square Root–Yeo Johnson transformation is performed to stabilize variance and make the data more normally distributed. **2. Build the time series pipeline** After preprocessing, the pipeline applies the **Vector Auto Regressor (VAR) algorithm**. VAR is a multivariate time series model that simultaneously forecasts multiple related variables. Unlike uni-variate models, VAR captures the influence of each metric on the others; for instance, how new subscriptions and cancellations jointly affect total revenue trends.The pipeline runs from the source dataset, through preprocessing, into the VAR algorithm, and finally outputs the results to the destination component. **3. Deploy the model via endpoint** After training, the model can be deployed directly from the Catalyst console by generating an **endpoint URL**. This endpoint allows external applications or dashboards to send recent billing data and receive real-time forecasts for future revenue and subscription metrics. Along with the predictions, the model also outputs evaluation metrics such as **MAPE, RMSE, MSE, MSLE, RMSLE, and SMAPE**, which help assess its accuracy. Also, the model uses cross validation metrics to test its robustness over time, and mean percentage error to ensure that the model’s forecasting performance remains consistent across different time windows. The final output comprises the forecasted daily values for key billing and subscription metrics, along with validation scores that measure model performance. These forecasts help in revenue planning, churn analysis, and promotional strategy optimization. -------------------------------------------------------------------------------- title: "Prerequisites" description: "Create a powerful ML pipeline that analyzed billing and subscription metrics for an OTT platform using the Catalyst QuickML components." last_updated: "2026-09-29T06:07:16.300Z" source: "https://docs.catalyst.zoho.com/en/tutorials/billing-and-subscription-metrics/prerequisites/" service: "All Services" related: - Time Series Algorithm (/en/quickml/help/ml-algorithms/time-series/) -------------------------------------------------------------------------------- # Prerequisites Because this tutorial is focused on **Catalyst QuickML**, all tasks—including dataset preparation, pipeline creation, model training, and deployment—will be performed within the Catalyst console. Before getting started, download the following dataset: [OTT_Platform_Billing_Subscription_Metrics.csv](https://workdrive.zoho.in/file/s3q013a8ee619ac214b9aad9f57092e7ee3b8) This dataset will be preprocessed and used to train a **multivariate time series forecasting** model with VAR in QuickML. -------------------------------------------------------------------------------- title: "Create a project" description: "Create a powerful ML pipeline that analyzed billing and subscription metrics for an OTT platform using the Catalyst QuickML components." last_updated: "2026-09-29T06:07:16.300Z" source: "https://docs.catalyst.zoho.com/en/tutorials/billing-and-subscription-metrics/create-a-project/" service: "All Services" related: - Catalyst Projects (/en/getting-started/catalyst-projects) -------------------------------------------------------------------------------- # Create a Project Let’s create a Catalyst project from the Catalyst console. 1. Log into the Catalyst console, and click **Create new Project**. 2. Enter the project’s name as **TimeSeriesForecasting** (or a name you wish to give for the project) in the pop-up window that appears. 3. Click the **Create** button. Your project will be created and opened. -------------------------------------------------------------------------------- title: "Upload the dataset" description: "Create a powerful ML pipeline that analyzed billing and subscription metrics for an OTT platform using the Catalyst QuickML components." last_updated: "2026-09-29T06:07:16.300Z" source: "https://docs.catalyst.zoho.com/en/tutorials/billing-and-subscription-metrics/upload-dataset/" service: "All Services" related: - Create Your First pipeline (/en/quickml/help/create-ml-pipeline) -------------------------------------------------------------------------------- # Upload the Dataset Let's begin by uploading the dataset in Catalyst QuickML using the available dataset **connectors**. 1. Navigate to the QuickML service in the Catalyst console and click **Start Exploring**. 2. Navigate to the Datasets component and click **Import Dataset**. A pop-up will be displayed. 3. In the Import Dataset window, navigate to Files and click **Upload File**. 4. Upload the Billing Subscription Metrics dataset that you have downloaded already. 5. Set the Quotes Type as Double Quotes (") and Escape Character as Backslash (\), and click **Next**. 6. The name of the dataset will be auto-populated based on the name of the uploaded file. Edit it if required, and click **Upload**. The dataset is now uploaded successfully. Note: The dataset contains missing values, and these gaps can affect the overall quality score. The score will improve once we address this during the pre-processing configuration step in the tutorial. The dataset will be displayed in the All Datasets section. You can click on the desired dataset name to view the dataset's details Once you click on the billing subscription dataset in the list, you'll be redirected to the Dataset Details page. The dataset's statistical **profile overview, sample preview**, and visualization chart is shown below. -------------------------------------------------------------------------------- title: "Create time series pipeline" description: "Create a powerful ML pipeline that analyzed billing and subscription metrics for an OTT platform using the Catalyst QuickML components." last_updated: "2026-09-29T06:07:16.300Z" source: "https://docs.catalyst.zoho.com/en/tutorials/billing-and-subscription-metrics/create-time-series-pipeline/" service: "All Services" related: - Time Series Algorithm (/en/quickml/help/ml-algorithms/time-series/) -------------------------------------------------------------------------------- # Create Time Series Pipeline Now that the dataset is uploaded, we’ll create a **time series pipeline** using the uploaded dataset. 1. Navigate to the *Pipeline* component in the left menu and click on the **Create Pipeline** button. 2. In the *Create Pipeline* window, select the **Forecasting** pipeline model. 3. Name the pipeline as billing_subscription_metrics, set the operation type to Multivariate Forecasting, and choose the billing_subscription_metrics dataset from the dataset drop-down. After selecting the dataset, specify the date as the Date/Timestamp column and total revenue as the Target column. Click **Create** to create the pipeline. You'll land on the Time Series smart mode pipeline builder as seen in the below image. This is a pre-built pipeline template with required operations to build a forecasting pipeline seamlessly in a short time. The smart mode pipeline builder contains two main stages: Preprocessing and Algorithm, in the pipeline template as shown above. Also, if you navigate to the *Profile* tab in the builder beside the sample preview, there’s an option to show the profile for sample data and whole data. The profile for the sample data is generated by default after configuring each operation in the pre-processing and algorithm stages, but in case you need a profile of the whole dataset during the pipeline building process, select **Whole Data** and click **Generate**. # Pipeline Configuration The first step in building any ML model is data preprocessing. In the Smart Builder, we have **predefined preprocessing operations** required for time series data. Below are the list of operations to be configured in the preprocessing stage for our multivariate forecasting model: - **Frequency**: Daily (1-day interval) - **Imputation**: Fill in missing values in the Target column - **Transformation**: Square Root Transform—Yeo Johnson Let’s go through each stage and understand why it’s necessary to configure the model. # Frequency Time series data must follow a consistent time interval. Frequency alignment ensures that all records are uniformly spaced. Here, select the interval as 1 Day because the dataset contains daily billing subscription metrics. This ensures that the model can learn from a continuous daily sequence without irregular gaps. Once the frequency is set, click **Save** and **Run** the frequency operation. # Imputation Real-world time series data often contains missing values. Imputation is used to fill in these gaps so the sequence remains consistent. Here, the target column, Total Revenue—along with all other columns with missing values—are selected with datatype Decimal, and missing values are filled in using the mean of the available values. This approach fills in missing values in a way that maintains the overall continuity and trend of the time series, without distorting the data from sudden spikes or drops. Once the imputation is handled, click **Save** and **Run** the imputation operation. # Transformation Transformation refers to applying statistical data transformation techniques to the data (such as log transformation or differencing) in order to stabilize variance or stationarity. In this pipeline, a Square Root–Yeo Johnson power transformation is applied to normalize the data distribution and reduce skewness. This helps the model handle non-linear relationships more effectively and improves overall forecasting accuracy. Once the transformation is configured, click **Save** and **Run** the transformation operation. # Stationarity Test Upon executing the transformation operation, you can find the Stationarity Test tab appearing next to the Preview and Profile tabs. Click the **Stationarity Test** tab to view the results of the stationarity analysis performed on the transformed dataset. This tab helps you evaluate whether each time series feature in your dataset is stationary or non-stationary using two statistical tests—KPSS (Kwiatkowski-Phillips-Schmidt-Shin) and ADF (Augmented Dickey–Fuller). These tests are essential to verify that the data has a stable mean and variance over time indicates that the time-series is stationary, which is a necessity for forecasting models. You can gain certain stationarity insights using the below options. - **Transformation stage overview**: Located at the top left, this dropdown (set to Overview by default) lets you choose which transformed feature or dataset view you want to analyze for stationarity. You can select Overview to see the results for all columns together, or pick a specific feature such as Total Revenue, New Subscriptions, or Cancelled Subscriptions, to view detailed stationarity results for that individual column. Depending on the transformations applied (for example, Square Root, Yeo-Johnson, Log, or Differencing), the dropdown allows you to explore how each transformation impacts the stationarity of the selected variable. - **After/Before transformation toggle**: Right next to the dropdown, you’ll see two options: Before transformation (shows results on the raw data) and After transformation (shows results on the data after applying your transformation steps). In the screen below, After transformation is selected, showing that all features have become stationary after processing. - **p-Value Significance Selector**: Located toward the top center of the panel, this option lets you choose the significance level used to determine whether a feature is stationary. You can switch between 1%, 5%, or 10% thresholds depending on how strict you want the stationarity test to be. **Note**: Changing this value updates the Result column (Stationary/Non-Stationary) in both KPSS and ADF test results accordingly. - 1%: Very strict, only very small p-values indicate stationarity. - 5%: Standard threshold, commonly used in statistical testing. - 10%: More lenient: useful when working with noisy or short time series. Learn about the significance of p-Value **here** - **Sample Data/Whole Data toggle**: Located on the upper right side, this toggle allows you to decide whether the stationarity tests should run on a subset of your dataset or on the entire dataset. In the displayed screen, Sample Data is selected, meaning the stationarity report is generated using a sampled portion of the data. - **Summary panel**: Shown on the left side of the Stationarity Test tab, this panel provides a quick summary of how many features are stationary or non-stationary based on each statistical test. **Total Columns** displays the total number of time-series features analyzed. **KPSS Test** shows the number of columns classified as stationary or non-stationary according to the KPSS test. ADF Test shows the same classification results according to the **ADF test**. The underlying assumption is that the series should be made stationary according to both tests to help ensure a well-performing model. - **Results table**: Located in the main area of the screen, this table lists the detailed test results for each column in the dataset. It contains the following: - **Features list**: The names of each time series feature tested (e.g., Total Revenue, New Subscriptions, etc.). - **KPSS p-Value and Result**: The p-value and outcome from the KPSS test; higher p-values generally indicate stronger evidence of stationarity. - **ADF p-Value and Result**: The p-value and outcome from the ADF test; lower p-values generally indicate stronger evidence of stationarity. Now that the preprocessing stage configuration is completed, the dataset is ready for training. In the next step, the VAR algorithm will use this processed data to model the target variable and generate the corresponding forecasts. # Algorithm configuration After preprocessing the dataset, the next step is to configure and apply the forecasting algorithm. In this tutorial, we use **VAR**, a robust multivariate time series algorithm that captures relationships among multiple interdependent variables. VAR simultaneously analyzes several related variables; for example, Total Revenue, New Subscriptions, Cancellations, and Refunds, to understand how changes in one metric influence the others over time. In our parameter configuration, **Max Lag** is set to a value between 0 and 5. This parameter determines how many previous time steps (lags) the model will use to predict future values. Selecting an appropriate lag value helps the model capture temporal dependencies effectively while preventing overfitting. Here, click the **Algorithm** stage, choose **VAR** from the drop-down and retain with the given parameter configuration. Then, click **Save**. Now, execute the pipeline by clicking the **Execute** button at the top-right corner of the pipeline builder page. This will redirect you to the pipeline details page shown below, which provides details about the executed pipeline with execution status. You can see that the pipeline execution is successful. Click **Execution Stats** to view more details about each stage of the model execution. The time series forecasting model is ready to analyze the billing subscription metrics for the OTT platform, and its details can be accessed under the Model field. (Click on the **billing_subscription_metrics_model** following the successful execution of the pipeline.) Upon clicking the model hyperlink **billing_subscription_metrics_model**, the model details page will be opened with necessary information about model details, evaluation metrics, cross validation metrics, and versioning details. # Model evaluation The Model Details page contains the evaluation metrics that help you understand the model performance during the training and validation process, cross-validation metrics, and as well as model versioning details. Model versions are used specifically to track the improvement or degradation of the performance and accuracy of the model upon frequent model training. As shown in the image below, you can view information about the model on the Model Details page. Learn more about model evaluation metrics **here**. -------------------------------------------------------------------------------- title: "Create an Endpoint" description: "Create a powerful ML pipeline that analyzed billing and subscription metrics for an OTT platform using the Catalyst QuickML components." last_updated: "2026-09-29T06:07:16.301Z" source: "https://docs.catalyst.zoho.com/en/tutorials/billing-and-subscription-metrics/create-endpoint/" service: "All Services" related: - Pipeline Endpoints (/en/quickml/help/pipeline-endpoints) -------------------------------------------------------------------------------- # Create Endpoint Now create an endpoint for the above model to allow **external applications** to interact with the model seamlessly and get model responses. 1. Navigate to the *Endpoints* component in the left menu and click **Create Endpoint**. 2. Provide a name for the endpoint in the Endpoint Name field; (we'll name it **billing subscription model**), select endpoint type as **ML Model**, and select the model pipeline name from the drop-down values of the *Choose Model* field. Click **Create Endpoint**. 3. Once the endpoint is created, you can view its details and test the deployed model directly from the Test the model section. In this view (as shown in the screenshot), the Request panel displays the input format expected by the model—typically the most recent timestamp or feature values required for generating the next forecast. For our use case, the model accepts the latest date value as input because the VAR model predicts the target variable based on the sequence of historical timestamps and related features processed during training. When you click **Get Result**, the model processes this input and generates the forecasted value, which appears in the Response panel on the right. The output includes a predicted value associated with the next forecasted date. This helps users validate that the model is returning predictions in the correct structure and for the correct time step before integrating the endpoint into business applications. 4. Click **Publish** and use the endpoint URL to integrate the created ML model with any other applications. 5. Upon publishing the endpoint, it will generate API details—which include Endpoint API URL, scope, headers, and more—used to access the model from external applications. Note: You can also check out **Endpoints Authentication** document to implement pipeline authentication. This ensures secured access to endpoints, the ML models, and datasets.