Qwen3-VL-4.3B-semantic
An expanded-vocabulary variant of Qwen/Qwen3-VL-4B-Instruct, designed for use as a text encoder in Krea 2.
The original 36 decoder layers and vision tower are retained. The main architectural modification is an expanded embed_tokens / tied lm_head vocabulary containing additional whole-word English tokens.
What changed
BaseThis modelTokenizer length151,936268,590Added word tokens—116,654New parameters—298,634,240Hidden size2,5602,560Decoder layers3636Vision towerOriginalOriginalMain modification—Expanded embed_tokens / tied lm_head
Existing token IDs are preserved. The added vocabulary consists of English words that were previously represented by more than one tokenizer subword.
The additional embedding rows contain semantic initialization rather than random initialization.
Semantic initialization
The added embeddings were constructed in several stages:
English lexical candidates were selected from the external lexical resources described below.
Initial semantic representations were obtained from ConceptNet Numberbatch.
WordNet-derived hypernym relationships were used to impose additional lexical structure.
The semantic representations were aligned to the original Qwen embedding space using a linear least-squares projection based on the mean embeddings of the corresponding original subword representations.
The resulting vectors were used to initialize the added Qwen vocabulary rows.
This procedure should be understood as semantic embedding expansion and initialization, not as full pretraining or continued pretraining of the Qwen3-VL transformer.
Data and external resources
This model incorporates information derived from several external lexical and semantic resources.
ConceptNet Numberbatch
ConceptNet Numberbatch 19.08 was used as the initial semantic representation for the added vocabulary.
Source:
ConceptNet
ConceptNet Numberbatch
Version: 19.08
The applicable ConceptNet / Numberbatch license and attribution requirements apply to the corresponding derived data.
WordNet / Open Multilingual WordNet
WordNet-derived lexical relations, including hypernym information, were used during semantic embedding construction.
The local lexical environment used the NLTK omw-1.4 resource where applicable.
The applicable WordNet / OMW licensing and attribution requirements apply to the corresponding derived data.
Wiktionary / Wiktextract
Wiktionary-derived lexical information was obtained through Wiktextract / Kaikki where applicable.
The exact dump or snapshot used to construct the released lexical data should be recorded for reproducibility.
Wiktionary-derived data is subject to the applicable Wiktionary licensing terms, including CC BY-SA and GFDL where applicable.
