core/oga - oga

Commit Graph

Author	SHA1	Message	Date
Yorick Peterse	74130df40d	XPath evaluation support for absolute paths.	2014-07-09 09:28:00 +02:00
Yorick Peterse	a398af19fe	Use a stack based XPath evaluator. This evaluator is capable of processing very, very basic XPath expressions. It does not yet know anything about axes, functions, operators, etc.	2014-07-08 23:57:21 +02:00
Yorick Peterse	9c661e1e60	Added XML::NodeSet#+ and XML::NodeSet#to_a	2014-07-08 23:25:09 +02:00
Yorick Peterse	266c66569e	Removed internal state of XPath::Evaluator. Come to think of it it might actually be easier to implement the evaluator as an actual VM. That is, instead of directly running on the AST it runs on some flavour of bytecode. Alternatively it runs directly on the AST but behaves more like a (stack based) VM. This would most likely be easier than passing a cursor to every node processing method.	2014-07-08 19:19:59 +02:00
Yorick Peterse	808e1e8c47	Initial, half-assed attempt at an XPath evaluator.	2014-07-07 19:41:09 +02:00
Yorick Peterse	8b381ac970	Added Node#remove. This method can be used to remove individual nodes without first having to retrieve the NodeSet they are stored in.	2014-07-04 10:26:41 +02:00
Yorick Peterse	e334e50ca6	Added Node#previous_element and Node#next_element. These methods can be used similar to #previous and #next expect that they only return Element instances opposed to all Node instances.	2014-07-04 10:18:18 +02:00
Yorick Peterse	94965961ce	Added XML::Element#inner_text. This method can be used to retrieve the text of the given node only. In other words, unlike Element#text it does not contain the text of any child nodes.	2014-07-04 09:50:47 +02:00
Yorick Peterse	a15c19c4af	Retrieving root nodes using Node#root_node. This method uses a loop to traverse upwards the DOM tree in order to find the root document/element. While this might have an impact on performance I don't expect Oga itself to call this method very often. The benefit is that Node instances don't require users to manually pass the top level document as an argument.	2014-07-03 00:26:15 +02:00
Yorick Peterse	e69fbc3ea7	Remove ownership when using NodeSet#delete.	2014-07-01 10:15:39 +02:00
Yorick Peterse	4e2b7f529d	Cleared up docs a bit of XML::Node.	2014-07-01 10:03:04 +02:00
Yorick Peterse	30845e3d65	Reliably remove nodes from owned sets. The combination of iterating over an array and removing values from it results in not all elements being removed. For example: numbers = [10, 20, 30] numbers.each do \|num\| numbers.delete(num) end numbers # => [20] As a result of this the NodeSet#remove method uses two iterations: 1. One iteration to retrieve all NodeSet instances to remove nodes from. 2. One iteration to actually remove the nodes. For the second iteration we iterate over the node sets and then over the nodes. This ensures that we always remove all nodes instead of leaving some behind.	2014-07-01 09:57:43 +02:00
Yorick Peterse	5073056831	Added require() for StringIO. This is needed since the lexer checks if the input is a StringIO instance.	2014-07-01 09:37:52 +02:00
Yorick Peterse	8314d24435	First pass at using the NodeSet class. The documentation still leaves a lot to be desired and so does the API. There also appears to be a problem where NodeSet#remove doesn't properly remove all nodes from a set. Outside of that we're making slow progress towards a proper DOM API.	2014-06-30 23:03:48 +02:00
Yorick Peterse	e71fe3d6fa	Indexing of NodeSet instances. Similar to arrays the NodeSet class now allows one to retrieve nodes for a given index.	2014-06-29 23:59:27 +02:00
Yorick Peterse	253575dc37	Basic docs for the NodeSet class.	2014-06-29 21:34:18 +02:00
Yorick Peterse	d2e74d8a0b	Specs for the NodeSet class.	2014-06-26 19:52:03 +02:00
Yorick Peterse	eb9d4fbccc	Changed NodeSet to behave more like an Array.	2014-06-26 09:37:54 +02:00
Yorick Peterse	a98f50b63b	NodeSet#push should not take node ownership.	2014-06-26 09:37:30 +02:00
Yorick Peterse	15fa7a2068	Remove explicit index tracking of NodeSet. Instead the various nodes can use NodeSet#index (aka Array#index) instead. This has a slight performance overhead on very large (millions) of nodes but should be fine in most other cases.	2014-06-25 09:41:58 +02:00
Yorick Peterse	884dbd9563	Rough sketch for a NodeSet class.	2014-06-24 19:06:45 +02:00
Yorick Peterse	b8cc6b5031	XPath::Parser#parse returns a XPath::Node instance	2014-06-23 20:23:08 +02:00
Yorick Peterse	03be9f241c	Docs on Racc operator precedence.	2014-06-23 10:36:42 +02:00
Yorick Peterse	96a7a40fdc	Disable code coverage for JRuby specific code.	2014-06-23 09:42:14 +02:00
Yorick Peterse	d2f15e37d0	Corrected XPath operator precedence. The previous commit didn't fully change the operator precedence according to the XPath 1.0 specification. Also thanks to @whitequark for clearing up a few things about Racc's operator precedence system.	2014-06-23 00:30:42 +02:00
Yorick Peterse	a440d3f003	Fixed XPath operator precedence. Apparently using multiple `left` rules with T_AND and T_OR being separate solves this problem. Riiiiight....	2014-06-23 00:15:43 +02:00
Yorick Peterse	6a2f4fa82d	Parsing support for more XPath operators. This still messes up some tests due to botched token precedence (by the looks of it).	2014-06-22 21:27:53 +02:00
Yorick Peterse	514c342cab	Lex wildcards as T_IDENT instead of T_STAR.	2014-06-20 20:37:34 +02:00
Yorick Peterse	45db337c76	Use Array#unshift for multiple xpath call args.	2014-06-17 20:12:25 +02:00
Yorick Peterse	b3ffc28cc7	Removed shift/reduce conflict in the xpath parser.	2014-06-17 20:09:44 +02:00
Yorick Peterse	497f57ccd2	Basic parser setup for XPath function calls.	2014-06-17 19:57:17 +02:00
Yorick Peterse	894de7f909	Lex all XPath expressions in a single machine. This allows literal values such as strings and numbers to be used as function arguments.	2014-06-17 19:56:57 +02:00
Yorick Peterse	bb7af98257	Updated used ASTs for all XPath parser specs.	2014-06-17 18:51:33 +02:00
Yorick Peterse	2298ef618b	Reworked handling of relative vs absolute XPaths.	2014-06-16 20:19:39 +02:00
Yorick Peterse	eba2d9954d	Support for parsing basic XPath expressions.	2014-06-12 00:20:46 +02:00
Yorick Peterse	70f3b7fa92	Lex XPath operators using individual tokens. Instead of lexing every operator as T_OP they now use individual tokens such as T_EQ and T_LT.	2014-06-09 23:35:54 +02:00
Yorick Peterse	7244e28eec	Corrected docs of the xpath lexer.	2014-06-04 19:33:46 +02:00
Yorick Peterse	1d2f9e6db6	Added T_STAR as an XPath parser token.	2014-06-02 09:27:25 +02:00
Yorick Peterse	e11b9ed32c	Basic XPath parser setup.	2014-06-01 23:02:28 +02:00
Yorick Peterse	54de2df0c7	Support for lexing XPath wildcard expressions. To support this we need to require whitespace around the "*" operator. This is not ideal but it will do for now.	2014-06-01 23:01:24 +02:00
Yorick Peterse	8dd8d7a519	Basic working XPath lexer. This doesn't lex everything of the XPath specification just yet and needs more tests.	2014-06-01 19:24:35 +02:00
Yorick Peterse	a50b76a2d8	Cleaned up XPath lexer boilerplate a bit.	2014-05-29 19:25:49 +02:00
Yorick Peterse	e0b07332d9	Boilerplate for the XPath lexer.	2014-05-29 19:25:49 +02:00
Yorick Peterse	be3f8fb494	Removed the on_newline XML lexer callback.	2014-05-29 14:21:48 +02:00
Yorick Peterse	ead5c71fee	Cleaned up the XML parser grammar. This resolves all shift/reduce and reduce/reduce conflicts that were previously present.	2014-05-29 01:37:19 +02:00
Yorick Peterse	49780e2b04	Fix for useless XML parser rules. Something tells me that using : and \| in your syntax might not be the best decision.	2014-05-28 21:36:06 +02:00
Yorick Peterse	28edc7726f	Rewind IO input upon resetting the lexer.	2014-05-26 00:33:20 +02:00
Yorick Peterse	629dcd3fe6	Support for IO inputs in the lexer. Using IO/StringIO objects one can parse large XML files without first having to read the entire file into memory. This can potentially save a lot of memory at the cost of a slightly slower runtime. For IO like instances the lexer will consume the input line by line. If a String is given it's consumed as a whole instead. A small side effect of reading the input line by line is that text such as "foo\nbar" will be lexed as two tokens instead of one. Fixes #19.	2014-05-26 00:30:39 +02:00
Yorick Peterse	6b9d65923a	Use a method for getting input in the XML lexer. Instead of directly accessing the `data` instance variable the C/Java code now uses the method `read_data`. This is part of one of the various steps required to allow Oga to read data from IO like instances. It also means I can freely change the name of the instance variable without also having to change the C/Java code.	2014-05-21 00:27:23 +02:00
Yorick Peterse	cd0f3380c4	Merge multiple CDATA tokens into a single token. The tokens T_CDATA_START, T_TEXT and T_CDATA_END have been merged together into T_CDATA.	2014-05-19 09:36:19 +02:00

1 2 3 4

177 Commits