GENERAL INFORMATION ------------------- 1. Title of dataset Semantic fluency responses in Spanish (L1) and English (L2) from Spanish primary-school children 2. Authors Name: María Paz Suárez-Coalla Institution: Language Psychology Laboratory, Department of Psychology, University of Oviedo, Oviedo, Spain Email: suarezpaz@uniovi.es ORCID: https://orcid.org/0000-0001-9772-2680 Name: Uxue Pérez-Litago Institution: Language Psychology Laboratory, Department of Psychology, University of Oviedo, Oviedo, Spain Email: perezuxue@uniovi.es ORCID: https://orcid.org/0000-0002-8317-8159 Name: Lucía García Castro Institution: University of Oviedo, Oviedo, Spain Name: Cristina Esteban Saster Institution: Department of Organization Engineering, University of Burgos, Burgos, Spain Email: ces1006@alu.ubu.es ORCID: https://orcid.org/0009-0000-8750-2043 Name: José Manuel Galán Institution: Department of Organization Engineering, University of Burgos, Burgos, Spain Email: jmgalan@ubu.es ORCID: https://orcid.org/0000-0003-3360-7602 3. Dataset description This dataset contains the primary semantic-fluency response sequences used in the analyses of the related article. The participants were 192 Spanish-L1 primary-school pupils in Grades 4-6 learning English as a foreign language. Each child was assigned two semantic categories in Spanish (L1) and two different categories in English (L2) in a crossed design. The categories were animals, fruits, body parts, and clothing. Participants had two minutes per task to write as many category members as possible. The dataset contains 724 participant-language-category tasks and 9,070 written responses. Response order is retained within each task. Original participant codes and the mapping between those codes and the neutral identifiers in this dataset are not deposited. 4. Keywords semantic fluency; child EFL learners; lexical-semantic access; bilingual lexical processing; Spanish; English; serial position; primary education 5. Funding Spanish Ministry of Science, Innovation and Universities, predoctoral grant FPU21/02740. 6. Geographic location of data collection An urban mainstream state-subsidised primary school in a city in northern Spain. The school and city are not identified in the public dataset. ACCESS INFORMATION ------------------ 1. Licence Creative Commons Attribution 4.0 International (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ 2. Dataset DOI To be assigned by RIUBU. 3. Related publication Submitted manuscript. METHODOLOGICAL INFORMATION -------------------------- 1. Participants The dataset comprises 192 pupils: Grade 4, n=57; Grade 5, n=72; and Grade 6, n=63. Recorded sex was 94 boys, 97 girls, and one value not reported. All participants were Spanish-L1 speakers receiving approximately three hours of English-as-a-foreign-language instruction per week. 2. Task and design Each child completed two categories in Spanish and two different categories in English. Category-language assignment was crossed across pupils so that a participant did not repeat the same category in both languages. Language order was counterbalanced. Responses were transcribed in production order. 3. Ethics Guardians provided written consent, and the study procedure was approved by the Research Ethics Committee of the Principality of Asturias. 4. Data structure semantic_fluency.csv is the canonical data file. It has one row per semantic- fluency task and retains the complete ordered response sequence in the response_sequence field. Individual responses are separated by commas inside that field; standard CSV quoting protects those internal commas. A separate token-level table is not included because it would contain no new primary observations. It can be reconstructed from semantic_fluency.csv by splitting response_sequence at commas while retaining the within-row order; the resulting position is the serial position used in word-level analyses. Neutral participant identifiers (P001-P192) replace the original participant codes. No re-identification key is included. Missing recorded sex is represented as not_reported. Variable definitions and allowed values are provided in codebook.csv. 5. Software and formats All files use UTF-8 encoding. The data and codebook are supplied as CSV files and can be read with R, Python, LibreOffice, Microsoft Excel, or a text editor. No specialist software is required. FILE OVERVIEW ------------- readme.txt Dataset documentation. semantic_fluency.csv Primary data used in the article analyses: 724 rows and 8 columns, with one row per participant-language-category task. codebook.csv Definitions, data types, allowed values, and units for the eight variables in semantic_fluency.csv. checksums.txt SHA-256 checksums for verifying the integrity of the three files above. VARIABLE SUMMARY ---------------- task_id Neutral unique task identifier. participant_id Neutral participant identifier. grade Primary-school grade (4, 5, or 6). sex Recorded sex (boy, girl, or not_reported). language Task language (Spanish or English). category Semantic category (animals, fruits, body_parts, or clothing). response_count Number of responses in the task sequence. response_sequence Comma-delimited responses in production order.