الفريق العربي للبرمجةأرشيف المنتديات · 2000 – 2023
نسخة أرشيفية للقراءة فقط — التسجيل والمشاركة مغلقان، والمحتوى محفوظ كما كان.

سؤال يا شباب ... بسرعة ؟؟

مغلق
بدأه mohammedsr في 20 مايو 2003 · 0 رد · 608 مشاهدة · في قسم تكنولوجيا Microsoft .NET العام
مشاركة: واتساب X فيسبوك تيليجرام
#1

سلامي للجميع ...

بدي كــــــــــــــــــــــــــــــــــــــــــــــود ببحث عن كلمات داخل ملفات اكس ام ال ...

I want search engine that search word in xml files ....

XML Document Indexing

Project purpose:

The project attempts to construct a simple XML search engine that allows user search for XML files by interested terms.

Implementation Strategy:

Inverted file structure is one of the most popular methods in text information retrieval. It is considered to be both efficient and accurate. To create inverted file structure for a collection of xml documents, the three major steps are described below:

Step I: Constructing Document-Term (D-T) Matrix:

A program has to be created to parse each document and count the frequency of each distinct word until all the documents have been parsed, the result is a Document-Term matrix.

D1

D2

D3

D4...

Dn

T1

2

3

5

0

5

T2

5

6

8

4

1

T3

5

8

0

10

3

...

2

5

0

4

3

Tn

5

3

4

9

0

Ideas for development:

· Use at least 30 XML file in your test cases

· Develop a program to parse the XML files and to extract each distinct word in each XML file. We want here for simplicity reasons just the elements and not the attributes to be processed.

· Use any parser and any programming language to achieve this step.

· Develop a help function to split a string into distinct words is created and remove all noise words (those like: the an is are be not then…).

· The frequency of each distinct document (weight) has to be dynamically generated

· To handle the parsed terms, use an XML file to store each distinct term and its addresses in the XML files. Example:

· The DTD could look like:

Step II: Retrieval

The D-T matrix is then stored into XML for retrieval purposes. A retrieval status value (RSV) is then calculated for each query according to how many query terms are used and the frequency those terms appear in a document. For example, if T3 is enter as the query term, since T3 appears in both D1, D2, D4, and Dn, these documents will be retrieved, listed from the highest frequency to the lowest one, indicating the extent of relevance.

Step III: Web User Interface

A user interface has to be developed by a scripting language. You are allowed to use your favorite language (ASP, JSP, PHP…). Users are allowed to enter as many key words as they want. For simplicity again let the search engine implement just OR search. It retrieves documents as long as they contain at least one of the keywords. Write a help function to parse the input string into distinct keywords, and then query them against the XML index you have created

هذا الموضوع مغلق.

مواضيع مشابهة