Aktibong Pag-aaral para sa Data Annotation: Gemini + Adala Guide

Adala Framework- Aktibong Pag-aaral para sa Data Annotation

Binabago ng aktibong pag-aaral kung paano tayo nagsasanay AI modelo sa pamamagitan ng matalinong pagpili ng pinakamahalagang data para sa anotasyon. Kapag ipinares sa makapangyarihang LLM gaya ng Google Gemini, lumilikha ito ng mahusay na mga pipeline ng anotasyon na nagpapababa ng manu-manong pagsisikap habang pinapanatili ang mataas na kalidad ng data.

Tinatalakay ng gabay na ito kung paano bumuo ng mga ganitong pipeline gamit ang Adala framework – isang makapangyarihan ngunit hindi gaanong nagagamit na tool para sa autonomous data labeling.

Magpapatupad kami ng pang-uri ng sintomas na medikal na gumagamit ng Gemini's mga kakayahan sa pamamagitan ng isang structured active learning workflow.

Pag-unawa sa Aktibong Pag-aaral para sa Data Annotation

Pag-unawa sa Aktibong Pag-aaral para sa Data Annotation

Tinutugunan ng aktibong pagkatuto ang pangunahing hamon sa pinangangasiwaang pagkatuto : ang pagkuha ng malalaking dami ng may label na datos. Sa halip na sapalarang pumili ng mga punto ng datos para sa anotasyon, tinutukoy ng mga algorithm ng aktibong pagkatuto ang mga pinaka-nakapagtuturong sample na higit na makakatulong sa pagpapabuti ng modelo.

Bakit mahalaga ang aktibong pag-aaral:

Binabawasan ang mga gastos sa anotasyon sa pamamagitan ng pagtutuon ng pagsisikap ng tao kung saan ito pinakamahalaga.
Nagpapabuti katumpakan ng modelo na may mas kaunting mga halimbawang may label.
Tinutugunan ang kawalan ng balanse ng klase sa pamamagitan ng pagbibigay-priyoridad sa mga kategoryang kulang sa representasyon.
Lumilikha ng tuluy-tuloy na pag-aaral sa pagitan ng modelo at annotator.

Dinadala ng balangkas ng Adala ang mga benepisyong ito sa mga daloy ng trabaho sa produksyon sa pamamagitan ng pagbibigay ng mga modular na bahagi na nagpapadali sa proseso ng aktibong pagkatuto . Bago tayo tumungo sa implementasyon, suriin muna natin kung bakit partikular na angkop ang Adala para sa integrasyon sa mga modernong LLM tulad ng Google Gemini.

Ano ang Adala? Isang Panimula sa Framework

Ang Adala (Autonomous Data Labeling Agent) ay isang open-source framework na sadyang idinisenyo para sa pagpapatupad ng mga espesyalisadong ahente para sa pagproseso ng datos . Hindi tulad ng mga tradisyunal na tool sa annotation, tinatanggap ng Adala ang isang agent-based na diskarte na pinagsasama ang:

Arkitekturang nakabatay sa kasanayan: Tukuyin ang mga partikular na kakayahan na kailangan ng iyong ahente ng anotasyon.
Flexibility ng runtime: Magpalit sa pagitan ng iba't ibang LLM o custom na runtime.
Mga koneksyon sa kapaligiran: Makipag-ugnayan sa iba't ibang pinagmumulan ng data.
Built-in na mga loop sa pag-aaral: Sanayin ang mga ahente upang mapabuti sa paglipas ng panahon.

Sa pagtingin sa halimbawa ng mabilisang pagsisimula ni Adala, makikita natin kung paano nito binubuo ang klasipikasyon ng damdamin :

python

import pandas as pd
from adala.agents import Agent
from adala.environments import StaticEnvironment
from adala.skills import ClassificationSkill
from adala.runtimes import OpenAIChatRuntime
from rich import print

# Train dataset
train_df = pd.DataFrame([
    ["It was the negative first impressions, and then it started working.", "Positive"],
    ["Not loud enough and doesn't turn on like it should.", "Negative"],
    ["I don't know what to say.", "Neutral"],
    ["Manager was rude, but the most important that mic shows very flat frequency response.", "Positive"],
    ["The phone doesn't seem to accept anything except CBR mp3s.", "Negative"],
    ["I tried it before, I bought this device for my son.", "Neutral"],
], columns=["text", "sentiment"])

# Test dataset
test_df = pd.DataFrame([
    "All three broke within two months of use.",
    "The device worked for a long time, can't say anything bad.",
    "Just a random line of text."
], columns=["text"])

agent = Agent(
    # connect to a dataset
    environment=StaticEnvironment(df=train_df),
    # define a skill
    skills=ClassificationSkill(
        name='sentiment',
        instructions="Label text as positive, negative or neutral.",
        labels=["Positive", "Negative", "Neutral"],
        input_template="Text: {text}",
        output_template="Sentiment: {sentiment}"
    ),
    # define runtimes
    runtimes = {
        'openai': OpenAIChatRuntime(model='gpt-4o'),
    },
    teacher_runtimes = {
        'default': OpenAIChatRuntime(model='gpt-4o'),
    },
    default_runtime='openai',
)

agent.learn(learning_iterations=3, accuracy_threshold=0.95)
predictions = agent.run(test_df)

Para sa aming gawain sa pag-uuri ng mga sintomas sa medisina, iaangkop namin ang arkitekturang ito upang maisama ang Google Gemini habang ipinapatupad ang isang pasadyang estratehiya sa aktibong pagkatuto.

Pag-set up ng Iyong Kapaligiran

Pabayaan's magsimula sa pamamagitan ng pag-install ng Adala at mga kinakailangang dependencies:

python

# Install Adala directly from GitHub
!pip install -q git+https://github.com/HumanSignal/Adala.git

# Verify installation
!pip list | grep adala

# Install additional dependencies
!pip install -q google-generativeai pandas matplotlib numpy

Kakailanganin din naming i-clone ang repository para sa direktang pag-access sa mga bahagi nito:

python

# Clone the repository for access to source files
!git clone https://github.com/HumanSignal/Adala.git

# Ensure the package is in our Python path
import sys
sys.path.append('./Adala')

# Import key components
from Adala.adala.annotators.base import BaseAnnotator
from Adala.adala.strategies.random_strategy import RandomStrategy
from Adala.adala.utils.custom_types import TextSample, LabeledSample

Pagsasama ng Google Gemini bilang Custom Annotator

Hindi tulad ng orihinal na implementasyon na gumamit ng isang pangunahing wrapper sa paligid ng Google Gemini, bubuo kami ng isang mas matatag na annotator na sumusunod sa mga pattern ng disenyo ng Adala. Ginagawa nitong mas madaling mapanatili at mapalawak ang aming solusyon.

Una, kailangan nating i-set up ang Google Generative AI kliyente:

python

import google.generativeai as genai
import os

# Set API key from environment or enter manually
GEMINI_API_KEY = os.getenv("GEMINI_API_KEY") or getpass("Enter your Gemini API Key: ")
genai.configure(api_key=GEMINI_API_KEY)

Ngayon, gagawa kami ng custom na annotator sa pamamagitan ng pagpapalawak ng Adala's BaseAnnotator klase:

python

import json
import re
from typing import List, Dict, Any, Optional

class GeminiAnnotator(BaseAnnotator):
    """Custom annotator using Google Gemini for medical symptom classification."""
    
    def __init__(self, 
                 model_name: str = "models/gemini-2.0-flash-lite", 
                 categories: List[str] = None,
                 temperature: float = 0.1):
        """Initialize the Gemini annotator.
        
        Args:
            model_name: The Gemini model to use
            categories: List of valid classification categories
            temperature: Controls randomness in generation (lower = more deterministic)
        """
        self.model = genai.GenerativeModel(
            model_name=model_name,
            generation_config={"temperature": temperature}
        )
        self.categories = categories or ["Cardiovascular", "Respiratory", 
                                         "Gastrointestinal", "Neurological"]
    
    def _build_prompt(self, text: str) -> str:
        """Create a structured prompt for the model.
        
        Args:
            text: The symptom text to classify
            
        Returns:
            A formatted prompt string
        """
        return f"""Classify this medical symptom into one of these categories:
        {', '.join(self.categories)}.
        
        Return JSON format: {{"category": "selected_category", 
        "confidence": 0.XX, "explanation": "brief_reason"}}
        
        SYMPTOM: {text}"""
    
    def _parse_response(self, response: str) -> Dict[str, Any]:
        """Extract structured data from model response.
        
        Args:
            response: Raw text response from Gemini
            
        Returns:
            Dictionary containing parsed fields
        """
        try:
            # Extract JSON from response even if surrounded by text
            json_match = re.search(r'(\{.*\})', response, re.DOTALL)
            result = json.loads(json_match.group(1) if json_match else response)
            return {
                "category": result.get("category", "Unknown"),
                "confidence": result.get("confidence", 0.0),
                "explanation": result.get("explanation", "")
            }
        except Exception as e:
            return {
                "category": "Unknown",
                "confidence": 0.0,
                "explanation": f"Error parsing response: {str(e)}"
            }
    
    def annotate(self, samples: List[TextSample]) -> List[LabeledSample]:
        """Annotate a batch of text samples.
        
        Args:
            samples: List of TextSample objects
            
        Returns:
            List of LabeledSample objects with annotations
        """
        results = []
        for sample in samples:
            prompt = self._build_prompt(sample.text)
            try:
                response = self.model.generate_content(prompt).text
                parsed = self._parse_response(response)
                
                # Create labeled sample with metadata
                labeled_sample = LabeledSample(
                    text=sample.text,
                    labels=parsed["category"],
                    metadata={
                        "confidence": parsed["confidence"],
                        "explanation": parsed["explanation"]
                    }
                )
            except Exception as e:
                # Graceful error handling
                labeled_sample = LabeledSample(
                    text=sample.text,
                    labels="Unknown",
                    metadata={"error": str(e)}
                )
            
            # Store reference to original sample
            labeled_sample._sample = sample
            results.append(labeled_sample)
            
        return results

Ang pagpapatupad na ito ay nagbibigay ng makabuluhang pagpapabuti sa orihinal:

  1. Ito ay sumusunod sa wastong class inheritance mula kay Adala's BaseAnnotator
  2. Nagpapatupad ng mga paraan ng pribadong helper para sa mabilis na pagbuo at pag-parse ng tugon
  3. Gumagamit ng structured maling paghawak at uri ng mga pahiwatig
  4. Nagbibigay ng kumpletong dokumentasyon

Pagbuo ng Pipeline ng Pag-uuri ng Sintomas

Gumawa tayo ng dataset ng mga sintomas na medikal para sa ating gawain sa pag-uuri. Hindi tulad ng orihinal na implementasyon, gagamit tayo ng mas magkakaibang dataset na may balanseng representasyon sa iba't ibang kategorya:

python

# Create a more comprehensive dataset
symptom_data = [
    # Cardiovascular symptoms
    "Chest pain radiating to left arm during exercise",
    "Heart palpitations when lying down",
    "Swollen ankles and shortness of breath",
    "Dizziness when standing up quickly",
    
    # Respiratory symptoms
    "Persistent dry cough with occasional wheezing",
    "Shortness of breath when climbing stairs",
    "Coughing up yellow or green mucus",
    "Rapid breathing with chest tightness",
    
    # Gastrointestinal symptoms
    "Stomach cramps and nausea after eating",
    "Burning sensation in upper abdomen",
    "Frequent loose stools with abdominal pain",
    "Yellowing of skin and eyes",
    
    # Neurological symptoms
    "Severe headache with sensitivity to light",
    "Numbness in fingers of right hand",
    "Memory loss and confusion",
    "Tremors in hands when reaching for objects"
]

# Convert to TextSample objects
text_samples = [TextSample(text=text) for text in symptom_data]

Pagpapatupad ng Advanced Active Learning Strategies

Ang orihinal na pagpapatupad ay gumamit ng isang simpleng mekanismo ng pagmamarka ng priyoridad. Papahusayin namin ito sa maraming diskarte para ipakita ang Adala's kakayahang umangkop:

python

import numpy as np
from typing import List, Callable

class PrioritizationStrategy:
    """Base class for sample prioritization strategies."""
    
    def score_samples(self, samples: List[TextSample]) -> np.ndarray:
        """Assign priority scores to samples.
        
        Args:
            samples: List of samples to score
            
        Returns:
            Array of scores, higher values indicate higher priority
        """
        raise NotImplementedError("Subclasses must implement this method")
    
    def select(self, samples: List[TextSample], n: int = 1) -> List[TextSample]:
        """Select the top n highest scoring samples.
        
        Args:
            samples: List of samples to select from
            n: Number of samples to select
            
        Returns:
            List of selected samples
        """
        if not samples:
            return []
        
        scores = self.score_samples(samples)
        indices = np.argsort(-scores)[:n]  # Descending order
        return [samples[i] for i in indices]

class KeywordPriority(PrioritizationStrategy):
    """Prioritize samples based on medical urgency keywords."""
    
    def __init__(self, keyword_weights: Dict[str, float]):
        """Initialize with keyword weights.
        
        Args:
            keyword_weights: Dictionary mapping keywords to priority weights
        """
        self.keyword_weights = keyword_weights
    
    def score_samples(self, samples: List[TextSample]) -> np.ndarray:
        scores = np.zeros(len(samples))
        for i, sample in enumerate(samples):
            # Base score
            scores[i] = 0.1
            
            # Add weights for each keyword found
            text_lower = sample.text.lower()
            for keyword, weight in self.keyword_weights.items():
                if keyword in text_lower:
                    scores[i] += weight
        
        return scores

class UncertaintyPriority(PrioritizationStrategy):
    """Prioritize samples based on model uncertainty."""
    
    def __init__(self, model_fn: Callable[[List[TextSample]], List[float]]):
        """Initialize with uncertainty model function.
        
        Args:
            model_fn: Function that returns uncertainty scores for samples
        """
        self.model_fn = model_fn
    
    def score_samples(self, samples: List[TextSample]) -> np.ndarray:
        # Higher uncertainty = higher priority
        return np.array(self.model_fn(samples))

# Create a combined strategy
keyword_weights = {
    "chest": 0.5,
    "pain": 0.4,
    "breathing": 0.4, 
    "dizz": 0.3,
    "head": 0.2,
    "numb": 0.2
}

keyword_strategy = KeywordPriority(keyword_weights)

Ngayon, hayaan's ipatupad ang aming pinahusay na aktibong loop sa pag-aaral:

python

from matplotlib import pyplot as plt
from IPython.display import clear_output
import time

def run_active_learning_loop(
    samples: List[TextSample],
    annotator: GeminiAnnotator,
    strategy: PrioritizationStrategy,
    iterations: int = 5,
    batch_size: int = 1,
    visualization_interval: int = 1
):
    """Run an active learning loop with visualization.
    
    Args:
        samples: Pool of unlabeled samples
        annotator: Annotation system
        strategy: Sample selection strategy
        iterations: Number of learning iterations
        batch_size: Samples to annotate per iteration
        visualization_interval: How often to update visualizations
    
    Returns:
        List of labeled samples
    """
    labeled_samples = []
    remaining_samples = list(samples)
    
    print("\nStarting Active Learning Loop:")
    
    for i in range(iterations):
        print(f"\n--- Iteration {i+1}/{iterations} ---")
        
        # Filter out already labeled samples
        remaining_samples = [
            s for s in remaining_samples 
            if s not in [getattr(l, '_sample', l) for l in labeled_samples]
        ]
        
        if not remaining_samples:
            print("No more samples to label. Stopping.")
            break
        
        # Select most important samples
        selected = strategy.select(remaining_samples, n=batch_size)
        
        # Annotate selected samples
        newly_labeled = annotator.annotate(selected)
        labeled_samples.extend(newly_labeled)
        
        # Display annotation results
        for sample in newly_labeled:
            print(f"Text: {sample.text}")
            print(f"Category: {sample.labels}")
            print(f"Confidence: {sample.metadata.get('confidence', 0):.2f}")
            explanation = sample.metadata.get('explanation', '')
            print(f"Explanation: {explanation[:100]}..." if len(explanation) > 100 else explanation)
            print()
        
        # Visualize results periodically
        if (i + 1) % visualization_interval == 0:
            visualize_results(labeled_samples)
            
    return labeled_samples

def visualize_results(labeled_samples: List[LabeledSample]):
    """Create visualizations of annotation results.
    
    Args:
        labeled_samples: List of labeled samples to visualize
    """
    if not labeled_samples:
        return
        
    # Extract data
    categories = [s.labels for s in labeled_samples]
    confidence = [s.metadata.get("confidence", 0) for s in labeled_samples]
    texts = [s.text[:30] + "..." for s in labeled_samples]
    
    # Set up plots
    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(15, 6))
    
    # Plot 1: Confidence by category
    category_counts = {}
    category_confidence = {}
    
    for cat, conf in zip(categories, confidence):
        if cat not in category_counts:
            category_counts[cat] = 0
            category_confidence[cat] = 0
        category_counts[cat] += 1
        category_confidence[cat] += conf
    
    for cat in category_confidence:
        category_confidence[cat] /= category_counts[cat]
    
    cats = list(category_counts.keys())
    counts = list(category_counts.values())
    avg_conf = list(category_confidence.values())
    
    x = np.arange(len(cats))
    width = 0.35
    
    ax1.bar(x - width/2, counts, width, label='Count')
    ax1.bar(x + width/2, avg_conf, width, label='Avg Confidence')
    ax1.set_xticks(x)
    ax1.set_xticklabels(cats, rotation=45)
    ax1.set_title('Category Distribution and Confidence')
    ax1.legend()
    
    # Plot 2: Individual sample confidence
    sorted_indices = np.argsort(confidence)
    ax2.barh(range(len(texts)), [confidence[i] for i in sorted_indices])
    ax2.set_yticks(range(len(texts)))
    ax2.set_yticklabels([texts[i] for i in sorted_indices])
    ax2.set_title('Sample Confidence')
    ax2.set_xlabel('Confidence')
    
    plt.tight_layout()
    plt.show()

Pagpapatakbo ng End-to-End Pipeline

Ngayon ay maaari na nating patakbuhin ang ating kumpletong pipeline ng aktibong pag-aaral:

python

# Initialize components
categories = ["Cardiovascular", "Respiratory", "Gastrointestinal", "Neurological"]
annotator = GeminiAnnotator(categories=categories)
strategy = keyword_strategy

# Run the active learning loop
labeled_data = run_active_learning_loop(
    samples=text_samples,
    annotator=annotator,
    strategy=strategy,
    iterations=5,
    visualization_interval=2
)

# Final visualization and analysis
visualize_results(labeled_data)

# Print summary statistics
print("\nAnnotation Summary:")
print(f"Total samples annotated: {len(labeled_data)}")

categories = [s.labels for s in labeled_data]
unique_categories = set(categories)
print(f"Categories found: {len(unique_categories)}")
for category in unique_categories:
    count = categories.count(category)
    print(f"  - {category}: {count} samples ({count/len(labeled_data):.1%})")

avg_confidence = sum(s.metadata.get("confidence", 0) for s in labeled_data) / len(labeled_data)
print(f"Average confidence: {avg_confidence:.2f}")

Mga Praktikal na Aplikasyon at Extension

Ang pipeline na ito ay may maraming praktikal na aplikasyon na lampas sa medikal na pag-uuri ng sintomas:

1. Pagmo-moderate ng Nilalaman

Isaayos ayon sa kahalagahan nilalamang iniulat ng gumagamit
Tumutok sa mga kategoryang may mataas na peligro
Iangkop ang mga limitasyon ng kumpiyansa batay sa ty ng nilalaman

2. Pagsusuri ng Feedback ng Customer

Kilalanin ang kagyat mga isyu sa customer
Kunin ang mga umuusbong na problema sa produkto
I-ruta ang feedback sa mga naaangkop na team

3. Pagproseso ng Dokumento ng Pagsubok sa Klinikal

I-extract ang mga ulat ng masamang kaganapan
Pag-uri-uriin mga kinalabasan na naiulat ng mga pasyente
Unahin ang mga signal ng kaligtasan

Maaari mong palawigin ang pagpapatupad na ito sa pamamagitan ng:

Pagdaragdag ng feedback loop para sa pagpapabuti ng annotator
Pagpapatupad ng iba't ibang estratehiya sa pagpili (diversity, clustering)
Paglikha ng isang web interface para sa human-in-the-loop validation
Pag-enable pag-uuri ng maraming label para sa mga kumplikadong sintomas

Konklusyon

Ang pagsasama ng Adala at Google Gemini ay nagbibigay ng isang makapangyarihang balangkas para sa pagbuo ng mga intelligent annotation pipeline. Sa pamamagitan ng paggamit ng mga aktibong estratehiya sa pag-aaral , maaari nating lubos na mabawasan ang kinakailangang manu-manong pagsisikap habang pinapanatili ang mataas na kalidad na mga annotation.

Ang mga modular na disenyo ng pattern na ipinakita sa tutorial na ito ay nagbibigay-daan para sa madaling pag-aangkop sa iba't ibang mga domain at mga gawain sa anotasyon.

Para sa mga interesado pang magsaliksik pa, ang Adala GitHub repository ay nag-aalok ng mga karagdagang halimbawa at dokumentasyon upang mapalawak ang mga konseptong ito sa mas kumplikadong mga senaryo ng anotasyon.

Mag-iwan ng Sagot

Hindi ilalathala ang iyong email address. Ang mga kinakailangang field ay minarkahan ng *.

Ginagamit ng site na ito ang Akismet upang mabawasan ang spam. Alamin kung paano pinoproseso ang data ng iyong komento.

Sumali sa Aimojo Tribo!

Sumali sa 76,000+ miyembro para sa insider tips bawat linggo! 
🎁 BONUS: Kunin ang aming $200"AI Mastery Toolkit” LIBRE kapag nag-sign up ka!

Nagte-trend AI Kagamitan
Dreamz.ai

Gumawa, makipag-chat, at tingnan ang iyong pribado AI kasintahan, na may mga video message na kasama sa bawat premium plan. Walang Sensor na 18+ AI kasamang chat para sa mga nasa hustong gulang na nagnanais ng ganap na kontrol.

Fantasybox

AI Paglikha ng Imahe at Video para sa mga Nasa Hustong Gulang na Walang Awtomatikong Pag-renew Paglikha ng nilalamang NSFW na may bayad kada kredito mula sa mga text prompt

Humata

Gawing Agarang, Maaaksyunang Katalinuhan ang Anumang Dokumento Ang AI Pinapagana na PDF at Document Analysis Assistant

Iris.ai

AI Knowledge Foundation na Nagpapabago sa Data ng Enterprise tungo sa Mapagkakatiwalaan at Nauusig na Katalinuhan Katumpakan AI para sa Reguladong R&D at Pamamahala ng Kaalaman sa Enterprise

Undermind

AI Pinapagana na Paghahanap ng Panitikan na Nakakahanap ng mga Papel na Hindi Nakikita ng Iba Ang research assistant na ginawa para sa mga siyentipiko, ng mga siyentipiko.

© Copyright 2023 - 2026 | Maging isang AI Pro | Ginawa gamit ang ♥