Full text
ViSketch-GPT: Collaborative Multi-Scale Feature Extraction for Sketch Recognition and Generation⋆ Giulio Federicoa,b,∗,Giuseppe Amatoa,Fabio Carraraa,Claudio Gennaroaand Marco Di Benedettoa aInstitute of Information Science and Technologies (ISTI-CNR), Via Giuseppe Moruzzi 1, Pisa, 56127, PI, Italy bUniversity of Pisa, Pisa, 56127, PI, Italy ARTICLE INFO Keywords: Sketch Generation,Recognition,Retrieval Multi-Scale Methodology Denoising Diffusion Probabilistic Model Natural language processing Vector Quantised-Variational AutoEncoder Transformer Signed Distance Field Deep Learning Machine Learning Artificial Intelligence ABSTRACT Understanding the nature of human sketches is challenging because of the wide variation in how they are created. Recognizing complex structural patterns improves both the accuracy in recognizing sketches and the fidelity of the generated sketches. In this work, we introduce ViSketch-GPT, a novel algorithm designed to address these challenges through a multi-scale context extraction approach. The model captures intricate details at multiple scales and combines them using an ensemble-like mechanism, where the extracted features work collaboratively to enhance the recognition and generation of key details crucial for classification and generation tasks. The effectiveness of ViSketch-GPT is validated through extensive experiments on the QuickDraw dataset. Our model establishes a new benchmark, significantly outperforming existing methods in both classification and generation tasks, with substantial improvements in accuracy and the fidelity of generated sketches. The proposed algorithm offers a robust framework for understanding complex structures by extracting features that collaborate to recognize intricate details, enhancing the understanding of structures like sketches and making it a versatile tool for various applications in computer vision and machine learning. Acknowledgement This work has received financial support by the Horizon Europe Research & Innovation Programme under Grant agreement N. 101092612 (Social and hUman ceNtered XR - SUN project) and by project "Italian Strengthening of ESFRI RI RESILIENCE" (ITSERR) funded by the European Union under the NextGenerationEU funding scheme (CUP:B53C22001770006). 1. Introduction Recognizing patterns of complex structures is fundamental for both the recognition and generation of visual content. However it becomes particularly challenging in the domain of human sketches where there is no single way to draw, but rather a variety of styles and representations for the same entities. Current approaches struggle to recognize and capture complex patterns, as they often fail to identify the intricate details that distinguish one entity from another, which are essential for accurate recognition and would enable better generation even when treating sketches as vector representations and leveraging NLP techniques to capture intricate temporal and structural dependencies. Vector sketches are inherently represented as ordered sequences of strokes, where each stroke is defined by a series of connected points. This sequential structure aligns naturally with NLP methodologies, which are designed to handle ordered, context-dependent data. Models like SketchRNN (Ha and Eck (2018)) have demonstrated the importance of leveraging sequential stroke information for sketch generation, employing recurrent architectures (RNNs) to predict the next stroke based on ⋆ ∗Corresponding author [email protected] (G. Federico); [email protected] (G. Amato); [email protected] (F. Carrara); [email protected] (C. Gennaro); [email protected] (M. Di Benedetto) ORCID(s): 0009-0005-0879-5631 (G. Federico); 0000-0003-0171-4315 (G. Amato); 0000-0001-5014-5089 (F. Carrara); 0000-0002-3715-149X (C. Gennaro); 0000-0001-5781-7060 (M. Di Benedetto) 1 Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 1 of 16 arXiv:2503.22374v1 [cs.CV] 28 Mar 2025
ViSketch-GPT previous ones. Building on this, Sketch-BERT (Lin, Fu, Xue and Jiang (2020)) extended the approach by adapting language modeling techniques such as BERT (Devlin, Chang, Lee and Toutanova (2019)) to the sketch domain. This enabled not only sketch generation but also improved recognition and retrieval, capitalizing on the transformer’s ability to capture bidirectional context. By treating sketches as sequences of visual "tokens," these approaches bridge the gap between textual and visual data, unlocking new possibilities for understanding and generating free-hand sketches. In this work, we present a novel framework that redefines how context is extracted and utilized for both recognition and generation. By decomposing the sketch into smaller patches using the quadtree technique and employing a multilevel context extraction mechanism for each patch, ViSketch-GPT captures contextual information at different scales, allowing each patch to be more accurately characterized. The collaborative integration of these features enables the model to capture intricate details, which are essential for precise recognition and significantly enhance the generation process by maintaining structural coherence. To evaluate the performance of ViSketch-GPT, we conducted experiments to assess its ability to recognize sketches and the fidelity of the generated sketches. Our results demonstrate that the model significantly outperforms the state of the art in classification, surpassing existing methods in both Top-1 and Top-3 accuracy, and demonstrating superior fidelity in sketch generation, as shown by the classifier’s ability to accurately recognize the new generated sketches. Our contributions are as follows: 1. We introduce a new methodology for capturing intricate details through the collaboration of multi-scale features, enhancing both sketch recognition and generation. 2. We evaluate the methodology through extensive experiments on the QuickDraw dataset, demonstrating its superior performance compared to state-of-the-art methods in both sketch recognition and generation. 3. We have optimized our approach to handling sparse data through a representation that mitigates potential issues during the generation phase. 2. Related works Since the introduction of Sketch-a-Net in 2015 (Yu, Yang, Song, Xiang and Hospedales (2015)) as a CNN-based model capable of generating free-hand sketches, significant advancements have been made in this domain in terms of architectures, representations, and datasets. In 2017, Google released QuickDraw, a large-scale sketch dataset comprising over 50 million sketches collected from players worldwide and SketchRNN (Ha and Eck (2018)) was introduced as an RNN-based deep Variational Autoencoder (VAE) capable of generating diverse sketches. Subsequent developments have focused on retrieval methods (Xu, Huang, Yuan, Pang, Song, Xiang, Hospedales, Ma and Guo (2018)), recognition approaches (Xu, Joshi and Bresson (2022); Hu, Li, Song, Xiang and Hospedales (2018)), and abstraction techniques aimed at simplifying sketches while preserving their recognizability (Muhammad, Yang, Song, Xiang and Hospedales (2018)), among others. Additionally, numerous new datasets have emerged, spanning both unimodal and multi-modal domains, as well as varying levels of granularity (coarsevs. fine-grained). Uni-modal datasets primarily support tasks such as recognition, retrieval, segmentation, and generation, whereas multi-modal datasets associate sketches with other modalities, including natural images, 3D models, and textual descriptions. Coarse-grained datasets (Eitz, Hays and Alexa (2012); Ha and Eck (2018)) provide more general sketch representations, while finegrained datasets offer higher levels of detail (Yu, Liu, Song, Xiang, Hospedales and Loy (2016)). Free-hand sketch tasks can be categorized into uni-modal and multi-modal tasks based on the type of data involved. Uni-modal tasks include recognition, retrieval, segmentation, and generation. Recognition aims to predict the class of a given sketch (Zhang, Liu, Zhang, Ren, Wang and Cao (2016a); Seddati, Dupont and Mahmoudi (2015); Zhang, Zhang and Qian (2016b); Lin et al. (2020); Guo, Wang, Roman-Rangel, Chao and Rui (2016); Ballester and Araujo (2016); Seddati, Dupont and Mahmoudi (2016); Zhang, She, Liu, Gan, Cao and Foroosh (2019); Jia, Fan, Yu, Liu, Wang and Latecki (2020)). Retrieval focuses on using a query sketch to retrieve similar samples from a dataset or collection (Lin et al. (2020); Wang and Li (2015); Xu et al. (2018); Creswell and Bharath (2016)). This is a particularly challenging task since traditional feature extraction methods (e.g., SIFT (Lowe (2004))) are ineffective due to the difficulty of identifying repeatable feature points across sketches drawn in diverse human styles. Segmentation involves the semantic partitioning of sketches and while conventional segmentation models for natural images could be adapted, the sequential nature of sketches has led to the development of dedicated models tailored to this task Wu, Qi, Liu and Yang (2018); Qi and Tan (2019); Wang, Lin, Wu, Li, Wang, Luo and He (2019); Kim, Wang, Öztireli and Gross (2018); Kaiyrbekov and Sezgin (2020); Yang, Zhuang, Fu, Wei, Zhou and Zheng (2021); Wang and Li (2024)). Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 2 of 16
ViSketch-GPT Deep learning-based approaches (Ha and Eck (2018); Cao, Yan, Shi and Chen (2019); Sasaki and Ogata (2018); Li, Gao, Shen, Zhang, Mei and Ren (2020); Ge, Goswami, Zitnick and Parikh (2020); Das, Yang, Hospedales, Xiang and Song (2020); Bhunia, Das, Muhammad, Yang, Hospedales, Xiang, Gryaditskaya and Song (2020); Ribeiro, Bui, Collomosse and Ponti (2020); Das, Yang, Hospedales, Xiang and Song (2021); Tiwari, Biswas and Lladós (2024)) have significantly outperformed traditional sketch generation methods. Sketch generation has numerous practical applications, including synthesizing new sketches, assisting artists in streamlining their design process, and reconstructing corrupted or incomplete sketches. A pioneering model in this domain is SketchRNN (Ha and Eck (2018)), which remains one of the state-of-the-art approaches for sketch-based abstraction and generalization. SketchRNN is a recurrent neural network-based generative model designed for both conditional and unconditional vector sketch generation. It was trained on QuickDraw, a largescale dataset of vector drawings collected from "Quick, Draw!", an online game where players were asked to sketch objects from a given category within 20 seconds. The dataset includes hundreds of object categories, each containing 70 training samples, along with 2.5K validation and test samples. The network is formally a Sequence-to-Sequence Variational Autoencoder (VAE), where the encoder is a bidirectional RNN that takes as input the sketch (the sequence defining it) and outputs a latent vector (the concatenation of the two hidden states obtained from the bidirectional RNNs). This latent vector is then transformed via a fully connected layer into a vector representing the mean and standard deviation, which are used to sample a latent vector. This sampled latent vector is provided as input to a decoder (an autoregressive RNN), which samples the subsequent strokes of the sketch. SketchBERT (Lin et al. (2020)) is a model based on BERT (Bidirectional Encoder Representations from Transformers Devlin et al. (2019)) adapted for handling free-hand sketches. It leverages BERT’s ability to capture contextual relationships between elements in a sequence to analyze and interpret sketches, treating them as temporal sequences of pen strokes. SketchBERT is designed for sketch recognition and generation tasks, utilizing a pre-trained representation to enhance performance across various computer vision and sketch generation tasks while maintaining high efficiency in the context of unstructured sequential data like sketches. AI-Sketcher (Cao et al. (2019)) proposes an enhancement of Sketch-RNN for handling multi-class generation, also based on a VAE generative model. They evaluated their method on a single dataset: FaceX. This dataset contains 5 million sketches of male and female facial expressions, and, unlike QuickDraw, the sketches were created by professionals. VASkeGAN (Balasubramanian, Balasubramanian et al. (2019)) combines a Variational Autoencoder (VAE) with a Generative Adversarial Network (GAN) to leverage the strengths of both models. The VAE is used to obtain an efficient representation of the data, while the GAN is employed to generate high-quality images. The goal is to produce visually appealing sketches, benefiting from both the compression capabilities of the VAE and the realistic generation abilities of the GAN. Additionally, a new metric called SkeScore was introduced; however, it is only applicable to vector-based generations. SketchGPT (Tiwari et al. (2024)) employs a sequence-to-sequence autoregressive model for sketch generation and completion by mapping complex sketches into simplified sequences of abstract primitives by leveraging the next token prediction objective strategy to understand sketch patterns, facilitating the creation and completion of drawings and also categorizing them accurately. 3. Methodology We formulate the problem as follows. We want to train a Neural Network that, given a specific class, is able to generate an image containing a sketch which belongs to the class. More formally, our objective is to learn the conditional distribution 𝑝(𝑆|𝑐)so that, given the class label 𝑐∈Z+, it is possible to generate 𝑆∈R𝐻×𝑊, representing a sketch of 𝑐. An example of such a task is shown in Figure 1. We split the generation process into two stages: i) pure generation and ii) refinement generation. In the pure generation stage, we tackle the task of modeling the distribution 𝑝(𝑆|𝑐)in a much smaller resolution space than the desired one. Specifically, if 𝐻, 𝑊 are the desired dimensions of the sketches, we choose 𝐻′, 𝑊 ′such that 𝐻′≪ 𝐻, 𝑊 ′≪ 𝑊 . More formally, this first stage can be defined as the process of inferring the conditional distribution: 𝑝(𝑆′|𝑐)where 𝑐∈Z+and 𝑆′∈R𝐻′×𝑊′ Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 3 of 16
ViSketch-GPT Figure 1: An example of the task we aim to tackle. Starting from the class label, we aim to generate a sketch belonging to that class. This simplifies and significantly accelerates the training process on complex shapes like sketches, as working in a lower-dimensional space reduces the complexity of modeling the distribution. Instead of accounting for all highly variable human-specific details, the model can focus on capturing the essential structural characteristics representative of the class, making learning more efficient. To model this distribution, we opted for diffusion model theory (Ho, Jain and Abbeel (2020)). In the generative refinement stage, our goal is to restore the details that were lost due to the low resolution of 𝑆′. For a single 𝑆′, there can exist multiple versions of 𝑆that contain different details, all statistically plausible. The objective of the second stage is to learn how to generate a plausible higher resolution version of 𝑆′, namely 𝑆, by inferring the following conditional distribution: 𝑝(𝑆|𝑆′, 𝑐)where 𝑐∈Z+, 𝑆 ∈R𝐻×𝑊and 𝑆′∈R𝐻′×𝑊′ An overview of the two stages is shown in Figure 2. Figure 2: Overview of the two stages: the first stage operates at a very low resolution to simplify and accelerate modeling; the second stage generates plausible details in a scalable manner. The way in which this distribution is modeled in the second stage is at the core of the algorithm. Further details are provided in the following sections. 3.1. Spatial Context Extraction for Scalable Refinement The goal of the generative refinement stage is to to model the distribution 𝑝(𝑆|𝑆′, 𝑐). In this way, it is possible to plausibly generate a higher resolution version consistent with 𝑆′(the output of the previous stage) and with the class 𝑐by generatively restoring its details. A naive approach would be to condition a generative network on 𝑆′and 𝑐to generate the entire 𝑆at its original resolution. Another approach involves dividing 𝑆into patches and performing super-resolution on the independent patches. These techniques may fail to fully capture the details we want to restore, let alone generate Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 4 of 16
ViSketch-GPT patches independently can lead to inconsistencies between adjacent patches previously generated, and generating a large number of patches—especially for sparse data—could be avoided to speed up the generation. The primary goal is therefore to model the distribution ensuring the coherent reconstruction of missing details while avoiding minimizing redundant generation. For this reason, we propose a novel pipeline. Given a ground-truth sketch 𝑆of the class 𝑐, and the low-resolution sketch 𝑆′, generated in Stage 1, the training process for the generative refinement performs preliminary steps to allow working effectively at patch level. The first step, performs a trivial resize operation on 𝑆′, to adapt it to the desired target resolution. We will refer to this resized version as 𝑆. Then, we build a quadtree by recursively partition 𝑆only if a certain tile contains significant information or until each leaf reaches the same resolution used in the first stage: 𝑊′×𝐻′(Figure 3). Note that when we build the quadtree, we ensure that all leaves, regardless of their level, have the same resolution ( 𝑙𝑠 ∈R𝑊′×𝐻′). Figure 3: First step of the generative refinement pipeline. Given the output of stage 1, 𝑆′is resized to the original resolution 𝑆and the quadtree is computed. The second step,copies the quadtree, obtained in previous step, onto the ground-truth sketch 𝑆(Figure 4). Figure 4: The second step of the generative refinement pipeline. Copy the quadtree of 𝑆into𝑆. We will thus model the distribution 𝑝(𝑆|𝑆′, 𝑐)as 𝑝(𝑆| 𝑆, 𝑐), or equivalently, in terms of the leaves of the quadtree: 𝑝(𝑆| 𝑆, 𝑐) = 𝑝({𝑙𝑠}|{ 𝑙𝑠}, 𝑐) which, by the property of joint probability, we can equivalently write as: Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 5 of 16
ViSketch-GPT 𝑝({𝑙𝑠}|{ 𝑙𝑠}, 𝑐)=𝑝(𝑙(1) 𝑠, ..., 𝑙(𝐿) 𝑠| 𝑙(1) 𝑠 , ..., 𝑙(𝐿) 𝑠 , 𝑐) =𝑝(𝑙(1) 𝑠|{ 𝑙𝑠}, 𝑐)⋅𝑝(𝑙(2) 𝑠|𝑙(1) 𝑠, 𝑙(2) 𝑠 , ..., 𝑙(𝐿) 𝑠 , 𝑐)... ... ⋅𝑝(𝑙(𝑖) 𝑠|𝑙(1) 𝑠, ..., 𝑙(𝑖−1) 𝑠, 𝑙(𝑖) 𝑠 , ..., 𝑙(𝐿) 𝑠 , 𝑐) .. ⋅𝑝(𝑙(𝐿) 𝑠|𝑙(1) 𝑠, ..., 𝑙(𝐿−1) 𝑠, 𝑙(𝐿) 𝑠 , 𝑐) = 𝐿 ∏ 𝑖=1 𝑝(𝑙(𝑖) 𝑠|𝑙(1) 𝑠, ..., 𝑙(𝑖−1) 𝑠, 𝑙(𝑖) 𝑠 , ..., 𝑙(𝐿) 𝑠 , 𝑐) (1) where 𝐿is the number of leaves. Each individual distribution is conditioned on the leaves of 𝑆but also on all the previously generated leaves of 𝑆, which will replace the corresponding ones in 𝑆. This will enable a coherent and aware generation of what has been previously generated. Note that the formulation (1) is quite naive and not scalable, as the generation of a single leaf needs all the initial and previously generated leaves. For a better scalability, to generate a particular leaf we will use a much lighter condition that we will indicate as the context of the leaf to generate. To this end, we define the context of a leaf 𝑙𝑠(𝑖)as the set of spatial neighbors of the leaf itself and all of its ancestors. Specifically, given the leaf 𝑙𝑠(𝑖), its spatial neighbors will correspond to the 3x3 grid of tiles where the leaf itself is at the center of the grid, and each tile has the same size as the leaf. The subsequent neighbors will be those of the 3x3 grid where the parent of the leaf is at the center, and each tile has the size of the parent. This process continues until the root itself is reached (the top-down view of 𝑆). The entire process is illustrated in Figure 5. This context satisfies three fundamental properties: •The context of a leaf 𝑙𝑠(𝑖)must contain a lossless information about the surrounding pixels. This will help generate a leaf that is consistent with the values of the surrounding pixels. •The context of a leaf 𝑙𝑠(𝑖)must contain progressively coarser information as it moves away from the target leaf. •The context of a leaf must provide an overview of the current version of 𝑆. Therefore, let 𝑆𝑖denote the data at step 𝑖of the refinement, i.e.: 𝑆𝑖= {𝑙(1) 𝑠, ..., 𝑙(𝑖−1) 𝑠, 𝑙(𝑖) 𝑠 , ..., 𝑙(𝐿) 𝑠 } and defining ℵ( 𝑙𝑠(𝑖))as the context of a leaf in 𝑆, the formulation (1) becomes: 𝑝({𝑙𝑠}|{ 𝑙𝑠}, 𝑐)= 𝐿 ∏ 𝑖=1 𝑝(𝑙(𝑖) 𝑠|ℵ( 𝑙𝑠(𝑖)), 𝑐)(2) During training, we will predict the generative refinement of the leaf 𝑙(𝑖) 𝑠using the class 𝑐and the context ℵ(⋅)of the leaf, as shown in Figure 6. 3.2. Generative refinement with autoregressive modeling To model each individual distribution in (2), we use a Transformer architecture (Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser and Polosukhin (2017)). Before training the Transformer 𝜃, we train a Vector Quantized Variational Autoencoder (VQ-VAE) (Van Den Oord, Vinyals et al. (2017)) to learn a discrete representation of the tiles at various resolutions they may have during the context creation process (Figure 5). This model acts as a tokenizer: each tile of size 𝑊′×𝐻′is encoded into a sequence of discrete indices belonging to a codebook. Once trained, the VQ-VAE is used to tokenize the refined patch into a sequence of discrete tokens, which are then used for training the Transformer. Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 6 of 16
ViSketch-GPT Figure 5: Process of creating the context of a leaf. Starting from the target leaf, the 3x3 tiles around it are taken with the leaf in the center. The same is done with the parent of the leaf until we reach the root itself. Each tile, regardless of the level, has the same resolution. Figure 6: To refine a leaf and make the approach more scalable, we consider its spatial context by including node’s adjacent neighborhood (i.e., centered 3x3 grid) for each level up to the root. The nearby nodes in blue, which extend beyond the image, are referred to as dummy nodes and have a fixed value. The encoder is a vision encoder (Dosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly et al. (2020)) that extracts features from 𝑆𝑖to enable the cross-attention mechanism with the decoder. However, unlike the classic ViT, we do not use all patches of 𝑆𝑖, as this would make the method unscalable. Instead, we select the context ℵ( 𝑙𝑠(𝑖))of the leaf to be refined. A notable property of our methodology is that the sequence produced by context extraction is always of the same length, regardless of the level at which the leaf is located. Indeed, if the octree has a maximum depth of 𝐷, then the context ℵ( 𝑙𝑠(𝑖))is a fixed sequence of length 𝐷× 9 + 1, where 9represents the neighbors at each level, and the addition of 1accounts for the contribution of the root. However, if a leaf belongs to an intermediate level 0< 𝑖 < 𝐷, its context still has the same length, but starting from a higher level, we use fixed and neutral values for the first elements of the sequence. The Transformer’s decoder outputs a sequence of probability distributions over the codebook indices, thus modeling the distribution of the refined leaf given the visual features extracted by the encoder. A visual representation of the architecture is shown in Figure 8. We train a VQ-VAE to compress each individual tile into a set of integer values 𝒛. To avoid low perplexity, and thus inefficient use of the codebook, we train the VQ-VAE not on the sparse data, but on a different representation that "intelligently fills" the empty spaces, preventing codebook collapse, which would introduce a strong bias in the transformer, leading to the prediction of repetitive sequences or overemphasizing certain tokens. This representation consists of calculating the Signed Distance Fields (SDF) of the sparse data. A SDF is a scalar field that represents Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 7 of 16
ViSketch-GPT the distance from a given point to the nearest surface. In our case, the surface corresponds to the stroke of the sketch. Positive values indicate points outside the stroke, while negative values indicate points inside it. This allows us to create a continuous representation of the shapes present in the sparse data, effectively filling in gaps and providing a more informative input for the VQ-VAE. Figure 7: Handling the Signed Distance Field (SDF) representation of sparse data helps VQ-VAE to have high perplexity (high codebook utilization) and thus avoid strong biases by the transformer in predicting certain classes. Specifically, the convolutional encoder performs a downsampling of each tile 𝑙to a smaller continuous spatial resolution: Encoder(𝑙)={𝑣1, 𝑣2, ...., 𝑣𝐾},where 𝑣𝑗∈𝑅𝐾,𝐷 Subsequently, each continuous encoding 𝑣𝑖will be mapped to the nearest element of the codebook of vectors P= {𝑝𝑖}𝑄 𝑖=1 ∈R𝑄,𝐷, where 𝑄is the total number of vectors in the codebook and 𝐷is the dimension of each vector. Therefore, in the end, each tile will be associated with a unique set of vectors: 𝐳= {𝑞1, 𝑞2, ..., 𝑞𝐾},where 𝑞𝑖= min 𝑝𝑗∈P||𝑣𝑖−𝑝𝑗|| whose codebook indices are chosen by the so-called quantizer Quantizer(𝑧). Each tile is then decompressed using a convolutional decoder: Decoder(𝐳) = 𝑙 The entire model is trained end-to-end by minimizing the following loss (Corona-Figueroa, Bond-Taylor, Bhowmik, Gaus, Breckon, Shum and Willcocks (2023)): 𝑉 𝑄 =𝑟𝑒𝑐 ( 𝑙, 𝑙)+𝑐𝑜𝑑𝑒𝑏𝑜𝑜𝑘 (𝑣, 𝑧) +𝑤𝑔⋅𝑔𝑒𝑛𝑒𝑟𝑎𝑡𝑜𝑟 ( 𝑙)+𝑤𝑑⋅𝑑𝑖𝑠𝑐𝑟𝑖𝑚𝑖𝑛𝑎𝑡𝑜𝑟 ( 𝑙, 𝑙) +𝑤𝑝⋅𝑝𝑒𝑟𝑐𝑒𝑝𝑡𝑢𝑎𝑙 ( 𝑙, 𝑙)(3) The entire pipeline (generation + generative refinement) is shown in Figure 8, while the pseudocode for training and inference is shown in algorithms (1) and (2). 4. Experimental Validation To validate our methodology, we tested ViSketch-GPT on two types of tasks: sketch generation and classification. 4.1. Dataset The proposed methodology is evaluated on the QuickDraw dataset (Ha and Eck (2018)), which is a collection of sketches created for the Google application Quick, Draw!, an online game in which users were asked to quickly draw, in less than 20 seconds, sketches related to specific categories. The dataset consists of 50 million sketches across 345 categories. Each sketch in QuickDraw is represented as a sequence of pen stroke actions, defined by five elements: (Δ𝑥,Δ𝑦, 𝑝1, 𝑝2, 𝑝3) where: Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 8 of 16
ViSketch-GPT Algorithm 1: Stage2: training 1repeat 2Choose random triplet (𝑆0∈R𝑊×𝐻, 𝑆′ 0∈R𝑊′×𝐻′, 𝑐 ∈Z+) 3Resize 𝑆′ 0to high resolution → 𝑆0∈R𝑊×𝐻 4Compute the leaves of 𝑆0via quadtree →{ 𝑙𝑠𝑖}𝐿 𝑖=1 5𝑙∼𝑈𝑛𝑖𝑓𝑜𝑟𝑚(1...𝐿) 6Replace the leaves preceding 𝑙with the corresponding tiles of 𝑆0 { 𝑙𝑠𝑖| 𝑙𝑠𝑖=𝑙𝑠𝑖for 𝑖= 1...(𝑙− 1)} 7Compute target leaf context →ℵ( 𝑙𝑠(𝑖)) 8Tokenize the refined leaf using the learned codebook 𝑡=Quantizer (Encoder(𝑙𝑠𝑖)) 9Compute the logits using Vision Encoder Decoder 𝐠=𝜃(ℵ( 𝑙𝑠(𝑖)), 𝑡, 𝑐) where 𝐠is a sequence of K logits of lenght Q. Each logit is a vector as long as the VQ-VAE codebook, containing information on which codebook index is most plausible to sample. 10 Apply softmax for each one of the K logits: 𝑦𝑘,𝑗 =𝑒𝑔𝑘,𝑗 ∑𝑄 𝑞′=1 𝑒𝑔𝑘,𝑞′ ,∀𝑗= 1,…, 𝑄 11 Compute the loss function (cross-entropy) = − 1 𝐾 𝐾 ∑ 𝑘=1 𝑄 ∑ 𝑞=1 𝑦𝑘,𝑞 log 𝑦𝑘,𝑞 12 Take gradient descent step on ∇𝜃 13 until convergence Figure 8: Architectural overview of the pure generation and generative refinement process. Given a class c, a low-resolution version of the SDF is generated, which is then scaled by a trivial resize to the desired resolution. The generative refinement stage, repeated for each leaf of the quadtree, restores the missing details. We end by clipping the SDF to get the data back in its sparse representation. •Δ𝑥,Δ𝑦represent the offset from the previous point. •𝑝1is a binary value indicating whether the pen is touching the paper (1) or lifted (0), thereby determining whether a line is drawn to the current point. •𝑝2is a binary indicator specifying whether the pen is lifted after reaching the current point (1) or not (0). Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 9 of 16
ViSketch-GPT Yu, Q., Liu, F., Song, Y.Z., Xiang, T., Hospedales, T.M., Loy, C.C., 2016. Sketch me that shoe, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Yu, Q., Yang, Y., Song, Y.Z., Xiang, T., Hospedales, T.M., 2015. Sketch-a-net that beats humans, in: British Machine Vision Conference. URL: https://api.semanticscholar.org/CorpusID:15004083. Zhang, H., Liu, S., Zhang, C., Ren, W., Wang, R., Cao, X., 2016a. Sketchnet: Sketch classification with web images, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Zhang, H., She, P., Liu, Y., Gan, J., Cao, X., Foroosh, H., 2019. Learning structural representations via dynamic object landmarks discovery for sketch recognition and retrieval. IEEE Transactions on Image Processing 28, 4486–4499. doi:10.1109/TIP.2019.2910398. Zhang, Y., Zhang, Y., Qian, X., 2016b. Deep neural networks for free-hand sketch recognition, in: Advances in Multimedia Information ProcessingPCM 2016: 17th Pacific-Rim Conference on Multimedia, Xi´ an, China, September 15-16, 2016, Proceedings, Part II, Springer. pp. 689–696. Giulio Federico, Giuseppe Amato, Fabio Carrara, Claudio Gennaro, Marco Di Benedetto: Preprint submitted to ElsevierPage 16 of 16